Sources & Licenses
Last updated: .
Shabdah is built mostly from original work — the flashcard structure, the illustration direction, the example curation, and the site itself. A few pieces draw on third-party open data, fonts, and software. This page lists everything we've identified, with the license that actually applies and where we sourced it, so it's separate from and doesn't get buried inside the Terms of Use.
A note on precision: we verified each license below against its primary source rather than assuming. Where a resource's licensing history is genuinely unclear even at the source, we've said so rather than guessing.
1. Shabdah original content
The flashcard set (word choices, translations, example selection, card layout and template, category structure), the site design, and the written content on this site are original to Picture the Words, © Shabdah, except where a third-party source is credited below.
Flashcard illustrations are produced with an AI image-generation tool (OpenAI's image API), directed and curated by Shabdah for each concept. They are not hand-drawn, and they are not stock photography or third-party artwork.
Hindi meanings and English translations for the curated concept set were initially machine-drafted, then reviewed and corrected by a native Hindi-speaking teacher before shipping. Concepts and fixes added after that review are verified instead using this project's own cross-language sense-checking process — comparing a candidate translation against the concept's category, its other-language translations, and its illustration — the same method used throughout this page's other sources. A small number of additional translations draw on an established third-party Hindi lexical resource (see §11), used only where neither path already supplied one.
2. Dictionary data — JMdict/EDICT
The Shabdah browser extension's word-lookup feature uses dictionary data from JMdict/EDICT, a Japanese–English dictionary file maintained by the Electronic Dictionary Research and Development Group (EDRDG).
- What we use it for: word lookup and definitions inside the browser extension. It is not bundled with the flashcard website's own data.
- License: Creative Commons Attribution-ShareAlike 4.0 International. Like all share-alike terms this requires that any adaptation we distribute is offered under the same licence to whoever receives it, not merely attributed — see Share-alike: what we offer back above.
- Commercial use: explicitly permitted by the license.
- Attribution / conditions: the license requires acknowledging the source in the product and linking to the project. It also asks that the data be kept reasonably up to date.
- Source & license: JMdict-EDICT Dictionary Project · EDRDG license terms
- Attribution: This product uses the JMdict/EDICT dictionary files, © the Electronic Dictionary Research and Development Group, used under a Creative Commons Attribution-ShareAlike 4.0 licence.
3. Example sentences — Tatoeba Project
Japanese example sentences shown on flashcards come from the Tatoeba Project, a free collaborative sentence database.
- What we use it for: example sentences displayed on the flashcard site.
- License: Creative Commons Attribution 2.0 France (CC BY 2.0 FR) by default. Tatoeba allows individual contributors to license their own sentences differently, so this default does not necessarily apply to every single sentence.
- Commercial use: generally permitted under CC BY; depends on the license chosen by the individual contributor for a given sentence.
- Attribution: required — CC BY requires crediting the author.
- Audio: Shabdah does not use Tatoeba's community audio recordings. All spoken audio on Shabdah is separately synthesized (see §9).
- Source & license: tatoeba.org · Tatoeba Terms of Use · CC BY 2.0 FR
4. Example sentences — Wiktionary (removed 2026-09-14)
A small number of flashcards briefly carried an example sentence from the English Wiktionary (via kaikki.org's structured extraction), used only where Tatoeba (§3) had no sentence for a word. Those 6,036 sentences have since been removed: Wiktionary is crowd-edited rather than native-speaker vetted, which is the standard the rest of this page holds to. 4,721 general-dictionary words and 59 of the 5,000 curated concepts now show no example sentence rather than an unvetted one. Kept here for the record, since the material was distributed for a period.
5. Parallel sentences — FLORES-101
The sentences that appear in more than two languages at once — the ones that let you read the same thought in Japanese, English, Hindi and French side by side — come from FLORES-101, Meta AI's evaluation benchmark. Its 2,009 sentences were translated from English into each language by professional translators and checked for quality, which is why they are the only sentences in Shabdah that carry all four languages on one row.
- What we use it for: 7,118 rows in the multilingual example corpus, covering the JA↔HI, JA↔FR and HI↔FR directions that no bilingual source reaches. Tatoeba (§3) is used first wherever it has a sentence.
- License: Creative Commons Attribution-ShareAlike 4.0 (CC BY-SA 4.0). Commercial use is explicitly permitted. Like all share-alike terms this requires that any adaptation we distribute is offered under the same licence to whoever receives it, not merely attributed.
- Attribution: required, and given here and in the credit line under every example block that uses it.
- Source & license: FLORES-101 · CC BY-SA 4.0
6. Hindi and French dictionaries — FreeDict and WOLF
Looking up a Hindi or French word that isn't one of Shabdah's own 5,000 concepts falls through to two open dictionaries, the way a Japanese word falls through to JMdict (§2).
- Hindi — FreeDict eng-hin. 25,642 English entries, built from the Shabdanjali dictionary released by IIIT Hyderabad. We invert it into Hindi → English, which is the direction a reader needs, and keep the example sentence each entry carries. 22,224 Hindi headwords result — see the quality note below.
- French — WOLF, the Wordnet Libre du Français. 55,373 French lemmas mapped onto Princeton WordNet synsets, which is where the English senses and the definitions come from. 55,356 headwords ship. A FreeDict fra-eng top-up used to be merged in on top of this and has been removed: it was GPL-licensed, it was worth 1,720 headwords out of 57,093, and dropping it leaves the French dictionary free of any copyleft obligation.
- Licenses: WOLF is CeCILL-C, a French free-software licence in the LGPL family. Princeton WordNet is under the permissive WordNet license. The Hindi dictionary is GNU GPL v2.0 or later — see “How we comply with the GPL” below.
- Commercial use: permitted by all three.
- Attribution: required; carried both on this page and inside the generated dictionary files themselves.
- Source & license: FreeDict · WOLF · Princeton WordNet
A note on WOLF's quality. Unlike the hand-compiled sources above, WOLF was built by automatically mapping Princeton WordNet onto French using bilingual dictionaries and parallel corpora, with partial rather than exhaustive manual review — so a French word's English sense is usually right and occasionally not, and no individual entry carries a human signature. It is used only for words outside Shabdah's own 5,000 concepts, which were translated and reviewed by people regardless.
A note on the Hindi fallback dictionary's quality. It is a volunteer-compiled dictionary (Shabdanjali, IIIT Hyderabad) of uneven quality by its own admission. We audited every entry rather than take that at face value: the bulk is sound (20,930 of 22,456 headwords are clean, with sensible glosses), and 230 unusable entries — a spreadsheet artifact carrying a mismatched gloss — plus a small number outside our content rules were removed. It is used only for words outside Shabdah's own 5,000 concepts, which were translated and reviewed by people regardless.
How we comply with the GPL. Shabdanjali was released under the GNU GPL by IIIT Hyderabad themselves, so this is not something FreeDict's packaging added and not something a different upstream would avoid. Rather than try to escape it, we comply with it. The GPL permits commercial use and permits charging money; what it asks is that the licensed file travels with its licence and its source.
Accordingly: hi-dictionary.json is distributed under the GNU GPL v2.0 or later
in its own right, carries that notice in its own _credit field, and the script
that generates it — tools/build_hi_fr_dictionaries.py, together with
tools/clean_gen_dictionaries.py — is its corresponding source and is
available on request from the contact page. The dictionary is a
separate data file fetched at runtime and is never compiled into the application, so it is an
aggregate alongside Shabdah rather than a part of it.
French no longer carries any GPL material at all.
A note on the phrase feature. The phrases shown on a card — faire du vélo, साइकिल चलाना, 自転車に乗る for “to cycle” — are not drawn from any dictionary on this page. They are Picture the Words' own concept labels, translated and reviewed by native speakers along with the rest of the concept set, simply collected together so you can see where one language needs several words for what another says in one. No new source, and nothing automatically generated.
This is a distinct feature from collocations, shown separately on the same card — see §12 below for how those are sourced.
7. Tanaka Corpus
Shabdah uses Tanaka Corpus data to help match example sentences to the word being studied (word indexing), rather than as the direct source of sentence text, which comes via Tatoeba (§3).
- Licensing status: unresolved. We checked two primary sources and they don't agree. Tatoeba's own downloads page describes the Tanaka Corpus material in its exports as belonging to the public domain. EDRDG's Tanaka Corpus documentation states the corpus was transitioned to CC BY 2.0 licensing in 2009, distinct from the original public-domain release by its creator, Professor Yasuhito Tanaka. We have not been able to establish which characterization governs the specific data Shabdah uses.
- What we're doing about it: pending clarification, we credit it as CC BY (the more restrictive of the two readings), since that's the safe assumption under either interpretation, and we extend the same share-alike offer to the data derived from it that we make for the sources that are unambiguously share-alike — see Share-alike: what we offer back. That is safe under every reading, including the public-domain one.
- Source references: Tatoeba downloads (Tanaka Corpus note) · EDRDG Tanaka Corpus history
Flagged for review: this licensing history should be confirmed (or a decision made on how conservatively to treat it) — see the note at the end of this page.
8. Fonts
All typefaces are loaded from Google Fonts under the SIL Open Font License 1.1, which permits commercial use and does not require attribution in the product itself.
- Fraunces — © The Fraunces Project Authors — OFL 1.1
- Inter — © The Inter Project Authors — OFL 1.1
- Noto Sans JP — © Adobe — OFL 1.1
- Noto Serif JP — © Google Inc. — OFL 1.1
9. Audio
Spoken audio for flashcards is synthesized using Microsoft Azure AI Speech (male English and Japanese neural voices). This is a paid commercial cloud service, not open community data, and no public attribution is required for using it in a downstream product. No community-recorded or crowdsourced audio (e.g. from Tatoeba) is used anywhere in the product.
10. Software libraries
- Supabase JS SDK (loaded via jsDelivr CDN) — © Supabase — MIT License. Commercial use permitted; the license notice is preserved by using the unmodified published package.
- kuromoji.js (bundled in the browser extension) — © 2014 Takuya Asano, © 2010–2014 Atilika Inc. and contributors — Apache License 2.0. The Japanese tokenizer that splits a sentence into words so hover lookup knows where one word ends and the next begins. Shipped unmodified, with its copyright and licence notices intact inside the file. Its own bundle includes zlib.js (© 2012 imaya, MIT) and async (MIT), whose notices travel with it.
- IPADIC (the dictionary kuromoji.js reads, bundled in the extension as
data/dict/) — © 2000–2007 Nara Institute of Science and Technology — IPADIC license, a BSD-style licence that permits commercial use and requires this notice be reproduced. Shipped unmodified. It supplies the part-of-speech and reading information behind Japanese word segmentation; it is not a source of any translation, definition or example sentence shown to a learner.
Payment processing is handled via Paddle and Razorpay's own checkout scripts, governed by their respective merchant terms of service rather than a content license, so they aren't listed as attributed content above.
11. Related words
The related-word chips shown on a card (synonyms, antonyms, and other closely associated concepts) are drawn from vetted lexical-relation sources per language, cross-checked against the concept's own category and its other-language translations before a candidate is accepted — the same adjudication standard applied throughout this project, not a bare dictionary lookup.
- Hindi: an established Hindi lexical-relation resource, used with the source's permission, and FreeDict eng-hin (§6), as two independent evidence channels.
- French: WOLF (§6), WoNeF, and FreeDict fra-eng.
- Japanese: JMdict (§2) and Japanese WordNet (NICT).
Licensing status: all sources listed here are confirmed permitted for this project's commercial use, each directly with the source where its own published license did not already make that clear.
12. Phrases & collocations
Distinct from the phrase feature described in §6 (concept labels combined, not drawn from a corpus), collocations are short, naturally-occurring word pairings extracted from real text and individually reviewed before shipping — concept by concept, not accepted automatically. Sourcing differs by language:
- Japanese: JMdict's own lexicographer-curated multi-word entries (§2) that contain the concept's headword — not extracted from a corpus.
- English & French: adjacency-frequency extraction from the Tatoeba sentence corpus (§3), topped up for French from a sample of French Wikipedia article text (wikimedia/wikipedia, CC BY-SA 3.0 + GFDL).
- Hindi: adjacency-frequency extraction from a sample of a large Hindi text corpus, used with the source's permission, topped up from a sample of Hindi Wikipedia article text (same source and licence as the French top-up above).
Nothing in this section is machine-translated or invented. Licensing status: all sources used in this section are confirmed permitted for this project's commercial use, each directly with the source where its own published license did not already make that clear. The Wikipedia-sourced portions are also share-alike licensed, so the derived collocation data built from them qualifies for the same redistribution offer already made for example-sentence data in “Share-alike: what we offer back” above.
13. Example sentences — AI-generated, audited corpus
Two kinds of example sentence appear on Shabdah, and this section exists so the difference is never left unstated. §3, §5 and §6 above are all human-sourced: real sentences written by people (Tatoeba contributors, FLORES-101's professional translators, FreeDict's lexicographers), never generated. The second kind is different, and is what this section covers: sentences generated by a large language model, used only where a human-sourced sentence was not available.
- Where these appear: two situations, and nowhere else. First, a concept in the core 5,000 whose target-language word Tatoeba and FLORES-101 have no matching sentence for. Second — the larger of the two by far — the example sentence shown for a specific inflected/conjugated form of a word (e.g. the negative, conditional, polite, or past-tense form). No corpus of real-world sentences pre-sorted by grammatical form like this exists to draw from, so every inflected-form example sentence on this site has always been generated rather than sourced, from the feature's introduction onward.
- How they were generated: by GPT and Gemini models, prompted with the concept's own verified English headword, category, and native-speaker-sourced target-language reference word or conjugated form (§1, and the dictionary/conjugation sources above) as the required sense to demonstrate — the model was never asked to invent a translation or a word, only a natural sentence using a translation that already existed and had already been vetted.
- What was never allowed to happen: this process does not touch meanings, translations, or the inflected forms themselves (the actual conjugated word, e.g. that a verb's past-negative-polite form is a specific string of characters) — all of that is native-speaker-sourced exactly as described elsewhere on this page and is never generated. Generation is confined to full example sentences built around an already-correct word.
- Auditing: LLM-generated text is not native-speaker-vetted by construction, so it was held to a higher bar before shipping, not a lower one. Every sentence in this corpus was individually reviewed by hand against its own text and the concept's reference data, with two independent AI systems (OpenAI, then Google Gemini) run first as automated flagging passes to surface candidate problems at this corpus's scale — never as the reviewer itself. Every flag either system raised was individually confirmed or rejected by hand before anything changed; nothing was removed, kept, or rewritten on an automated system's say-so alone. A confirmed-bad sentence is deleted outright, never patched with invented replacement text, since this corpus (unlike every other source on this page) has no vetted original to correct against. Where this leaves a concept or a specific inflected form with no example sentence at all, that gap is left honest rather than papered over.
- Ownership: generated output from these providers' APIs, used under their respective commercial API terms; Shabdah holds the rights to redistribute it as part of the product the same way it does any other generated asset (illustrations, synthesized audio, §9).
A note on legal certainty
This page reflects attribution and licensing requirements we identified from the available source licenses at the time of writing, verified against each resource's own primary documentation where possible. It is not a substitute for legal advice. One item above has a licensing question we could not fully resolve from primary sources: the Tanaka Corpus's licensing history (§7). The commercial-use permission questions previously open for the Hindi-language sources referenced in §11 and §12 have since been confirmed directly with those sources. If you have information that clarifies the remaining item, or a concern about anything on this page, please get in touch.