My Notes

Shabdah

Sources & Licenses

Last updated: .

Shabdah is built mostly from original work — the flashcard structure, the illustration direction, the example curation, and the site itself. A few pieces draw on third-party open data, fonts, and software. This page lists everything we've identified, with the license that actually applies and where we sourced it, so it's separate from and doesn't get buried inside the Terms of Use.

A note on precision: we verified each license below against its primary source rather than assuming. Where a resource's licensing history is genuinely unclear even at the source, we've said so rather than guessing.

Share-alike: what we offer back. Three of the sources below — JMdict/EDICT (§2), FLORES-101 (§5) and, on the more conservative reading of its history, the Tanaka Corpus (§7) — are licensed under Creative Commons Attribution-ShareAlike. Share-alike asks more than credit: where we distribute an adaptation of that material, the adaptation has to reach you under the same licence.

We do distribute adaptations. Our example-sentence data files (examples.json, the per-language examples-*.json files and sentence-search-index.json) are derived from those corpora — filtered, re-aligned to our concepts, and re-indexed. Those derived datasets are offered under the Creative Commons Attribution-ShareAlike 4.0 International licence, and you may redistribute and adapt them on those terms. Request a copy from the contact page.

This is a carve-out from the general restriction on redistributing Shabdah content in the Terms of Use, and the Terms say so. It covers the openly-licensed source data and material derived from it. It does not extend to the illustrations, the synthesized audio, the curated 5,000-concept set and its translations, the phrase and related-word data, or the site and extension code, none of which are derived from these corpora.

1. Shabdah original content

The flashcard set (word choices, translations, example selection, card layout and template, category structure), the site design, and the written content on this site are original to Picture the Words, © Shabdah, except where a third-party source is credited below.

Flashcard illustrations are produced with an AI image-generation tool (OpenAI's image API), directed and curated by Shabdah for each concept. They are not hand-drawn, and they are not stock photography or third-party artwork.

Hindi meanings and English translations for the curated concept set were initially machine-drafted, then reviewed and corrected by a native Hindi-speaking teacher before shipping. Concepts and fixes added after that review are verified instead using this project's own cross-language sense-checking process — comparing a candidate translation against the concept's category, its other-language translations, and its illustration — the same method used throughout this page's other sources. A small number of additional translations draw on an established third-party Hindi lexical resource (see §11), used only where neither path already supplied one.

2. Dictionary data — JMdict/EDICT

The Shabdah browser extension's word-lookup feature uses dictionary data from JMdict/EDICT, a Japanese–English dictionary file maintained by the Electronic Dictionary Research and Development Group (EDRDG).

3. Example sentences — Tatoeba Project

Japanese example sentences shown on flashcards come from the Tatoeba Project, a free collaborative sentence database.

4. Example sentences — Wiktionary (removed 2026-09-14)

A small number of flashcards briefly carried an example sentence from the English Wiktionary (via kaikki.org's structured extraction), used only where Tatoeba (§3) had no sentence for a word. Those 6,036 sentences have since been removed: Wiktionary is crowd-edited rather than native-speaker vetted, which is the standard the rest of this page holds to. 4,721 general-dictionary words and 59 of the 5,000 curated concepts now show no example sentence rather than an unvetted one. Kept here for the record, since the material was distributed for a period.

5. Parallel sentences — FLORES-101

The sentences that appear in more than two languages at once — the ones that let you read the same thought in Japanese, English, Hindi and French side by side — come from FLORES-101, Meta AI's evaluation benchmark. Its 2,009 sentences were translated from English into each language by professional translators and checked for quality, which is why they are the only sentences in Shabdah that carry all four languages on one row.

6. Hindi and French dictionaries — FreeDict and WOLF

Looking up a Hindi or French word that isn't one of Shabdah's own 5,000 concepts falls through to two open dictionaries, the way a Japanese word falls through to JMdict (§2).

A note on WOLF's quality. Unlike the hand-compiled sources above, WOLF was built by automatically mapping Princeton WordNet onto French using bilingual dictionaries and parallel corpora, with partial rather than exhaustive manual review — so a French word's English sense is usually right and occasionally not, and no individual entry carries a human signature. It is used only for words outside Shabdah's own 5,000 concepts, which were translated and reviewed by people regardless.

A note on the Hindi fallback dictionary's quality. It is a volunteer-compiled dictionary (Shabdanjali, IIIT Hyderabad) of uneven quality by its own admission. We audited every entry rather than take that at face value: the bulk is sound (20,930 of 22,456 headwords are clean, with sensible glosses), and 230 unusable entries — a spreadsheet artifact carrying a mismatched gloss — plus a small number outside our content rules were removed. It is used only for words outside Shabdah's own 5,000 concepts, which were translated and reviewed by people regardless.

How we comply with the GPL. Shabdanjali was released under the GNU GPL by IIIT Hyderabad themselves, so this is not something FreeDict's packaging added and not something a different upstream would avoid. Rather than try to escape it, we comply with it. The GPL permits commercial use and permits charging money; what it asks is that the licensed file travels with its licence and its source.

Accordingly: hi-dictionary.json is distributed under the GNU GPL v2.0 or later in its own right, carries that notice in its own _credit field, and the script that generates it — tools/build_hi_fr_dictionaries.py, together with tools/clean_gen_dictionaries.py — is its corresponding source and is available on request from the contact page. The dictionary is a separate data file fetched at runtime and is never compiled into the application, so it is an aggregate alongside Shabdah rather than a part of it.

French no longer carries any GPL material at all.

A note on the phrase feature. The phrases shown on a card — faire du vélo, साइकिल चलाना, 自転車に乗る for “to cycle” — are not drawn from any dictionary on this page. They are Picture the Words' own concept labels, translated and reviewed by native speakers along with the rest of the concept set, simply collected together so you can see where one language needs several words for what another says in one. No new source, and nothing automatically generated.

This is a distinct feature from collocations, shown separately on the same card — see §12 below for how those are sourced.

7. Tanaka Corpus

Shabdah uses Tanaka Corpus data to help match example sentences to the word being studied (word indexing), rather than as the direct source of sentence text, which comes via Tatoeba (§3).

Flagged for review: this licensing history should be confirmed (or a decision made on how conservatively to treat it) — see the note at the end of this page.

8. Fonts

All typefaces are loaded from Google Fonts under the SIL Open Font License 1.1, which permits commercial use and does not require attribution in the product itself.

9. Audio

Spoken audio for flashcards is synthesized using Microsoft Azure AI Speech (male English and Japanese neural voices). This is a paid commercial cloud service, not open community data, and no public attribution is required for using it in a downstream product. No community-recorded or crowdsourced audio (e.g. from Tatoeba) is used anywhere in the product.

10. Software libraries

Payment processing is handled via Paddle and Razorpay's own checkout scripts, governed by their respective merchant terms of service rather than a content license, so they aren't listed as attributed content above.

11. Related words

The related-word chips shown on a card (synonyms, antonyms, and other closely associated concepts) are drawn from vetted lexical-relation sources per language, cross-checked against the concept's own category and its other-language translations before a candidate is accepted — the same adjudication standard applied throughout this project, not a bare dictionary lookup.

Licensing status: all sources listed here are confirmed permitted for this project's commercial use, each directly with the source where its own published license did not already make that clear.

12. Phrases & collocations

Distinct from the phrase feature described in §6 (concept labels combined, not drawn from a corpus), collocations are short, naturally-occurring word pairings extracted from real text and individually reviewed before shipping — concept by concept, not accepted automatically. Sourcing differs by language:

Nothing in this section is machine-translated or invented. Licensing status: all sources used in this section are confirmed permitted for this project's commercial use, each directly with the source where its own published license did not already make that clear. The Wikipedia-sourced portions are also share-alike licensed, so the derived collocation data built from them qualifies for the same redistribution offer already made for example-sentence data in “Share-alike: what we offer back” above.

13. Example sentences — AI-generated, audited corpus

Two kinds of example sentence appear on Shabdah, and this section exists so the difference is never left unstated. §3, §5 and §6 above are all human-sourced: real sentences written by people (Tatoeba contributors, FLORES-101's professional translators, FreeDict's lexicographers), never generated. The second kind is different, and is what this section covers: sentences generated by a large language model, used only where a human-sourced sentence was not available.

A note on legal certainty

This page reflects attribution and licensing requirements we identified from the available source licenses at the time of writing, verified against each resource's own primary documentation where possible. It is not a substitute for legal advice. One item above has a licensing question we could not fully resolve from primary sources: the Tanaka Corpus's licensing history (§7). The commercial-use permission questions previously open for the Hindi-language sources referenced in §11 and §12 have since been confirmed directly with those sources. If you have information that clarifies the remaining item, or a concern about anything on this page, please get in touch.