Where Shirabe gets its data
Source versions
Every dataset we currently have loaded, with its origin and exact upstream build. Re-importing a source updates its row here, so this stays an accurate reference. Licences and fuller attribution follow below.
| Source | Origin | Licence | Version | Data date |
|---|---|---|---|---|
| JMdict: word entries | jmdict-simplified (EDRDG) | EDRDG | 3.6.2 | 2026-07-27 |
| JMnedict: names | jmdict-simplified (EDRDG) | EDRDG | 3.6.2 | 2026-07-27 |
| KANJIDIC2: kanji | jmdict-simplified (EDRDG) | EDRDG | 3.6.2 | 2026-07-27 |
| KRADFILE: kanji → components | jmdict-simplified (EDRDG) | CC BY-SA 4.0 | Not available | 2026-08-03 |
| RADKFILE: radical → kanji | jmdict-simplified (EDRDG) | CC BY-SA 4.0 | 3.6.2 | 2026-08-03 |
| Example sentences (Tatoeba) | Tatoeba (sense links: jmdict-simplified) | CC BY 2.0 FR | Not available | 2026-08-02 |
| Kanken levels (漢検) | mimneko/kanji-data | Not available | Not available | 2026-06-26 |
| Jōyō table (常用漢字表 本表) | mimneko/kanji-data | CC0 1.0 | Not available | 2026-02-10 |
| Jōyō appendix (付表) | mimneko/kanji-data | CC0 1.0 | Not available | 2026-02-09 |
| Radical names: Kanji alive | kanjialive | CC BY 4.0 | Not available | 2026-07-29 |
| Pitch accents: UniDic | UniDic (NINJAL) | Modified BSD | 202512 | 2025-12-31 |
| Word frequency: jiten-global | jiten.moe | CC BY-SA 4.0 | Not available | 2026-08-02 |
| Wikipedia abstracts | DBpedia | CC BY-SA 3.0 | 2016-10 | 2026-07-26 |
| BabelStone IDS: kanji structure | BabelStone (Andrew West) | Not available | Not available | 2025-06-27 |
| Stroke order: KanjiVG | KanjiVG (Ulrich Apel) | CC BY-SA 3.0 | Not available | 2025-08-16 |
Dictionary data
JMdict: Japanese↔multilingual word entries
Compiled by the Electronic Dictionary Research and Development Group (EDRDG, James Breen et al.) and distributed under the EDRDG licence. We use the jmdict-simplified JSON conversion by Stanislav Petrov.
JMnedict: proper-noun dictionary
Also from EDRDG, under the same licence terms, and consumed through the jmdict-simplified JMnedict release.
KANJIDIC2: kanji dictionary
EDRDG, under the same licence. Shirabe uses jmdict-simplified’s KANJIDIC2 all release, with kanji meanings in English, French, Spanish, and Portuguese.
JLPT levels (N5–N1)
The levels shown on kanji and words come from Jonathan Waller’s JLPT Resources, licensed CC BY 4.0. Because the post-2010 JLPT publishes no official kanji or vocabulary list, these are unofficial lists: kanji from the KANJIDIC snapshot and words through Bluskyo/JLPT_Vocabulary.
KRADFILE: kanji component breakdown
Kanji parts come from EDRDG’s KRADFILE / KRADFILE2 (Michael Raine, James Breen et al.), licensed CC BY-SA 4.0 and consumed through jmdict-simplified.
BabelStone IDS: kanji structure
Positional Ideographic Description Sequences (⿰⿱⿴…) come from Andrew West’s BabelStone IDS data, released to the public domain.
Kanken (漢字検定) levels
Assigned Kanji Kentei levels come from mimneko/kanji-data, released under CC0 1.0.
Jōyō kanji table (常用漢字表)
Official readings and examples come from the Japanese Agency for Cultural Affairs’ 2010 常用漢字表, digitised by mimneko/kanji-data under CC0 1.0. The underlying table is a Japanese cabinet notification and is not subject to copyright.
成り立ち: kanji formation and glyph origin
Formation types and Japanese glyph-origin notes come from Japanese Wiktionary under CC BY-SA 4.0. Shinjitai inherit their kyūjitai’s origin. A few jōyō gaps have an AI-assigned classification and are marked on the kanji page.
RADKFILE: search radicals to kanji
The 253 search radicals and their kanji come from EDRDG’s RADKFILE / RADKFILE2, licensed CC BY-SA 4.0 and consumed through jmdict-simplified.
Kangxi radical names and meanings
Japanese readings, English glosses, stroke counts, and positional categories come from Kanji alive, licensed CC BY 4.0.
Sixty variant forms (きへん, たつへん, あなかんむり and the rest) have no Unicode character at all, so the radical chart sets them in Kanji alive’s own Japanese Radicals font, derived from Source Han Sans and licensed Apache 2.0 by Adobe Systems. The font’s licence ships beside it in the repository.
Wikipedia abstracts via DBpedia
Lead summaries come from DBpedia’s long_abstracts dataset. The text belongs to its Wikipedia contributors and is dual-licensed under CC BY-SA 3.0 and the GNU Free Documentation License.
Example sentences: Tatoeba
Example sentences come from Tatoeba through the JMdict examples set, licensed CC BY 2.0 FR and consumed through jmdict-simplified.
Pitch accents: UniDic
Pitch-accent data comes from UniDic by the National Institute for Japanese Language and Linguistics, available under GPL 2.0, LGPL 2.1, or Modified BSD.
Word frequency: jiten.moe
Rank ordering for the frequency sort comes from the global frequency list published by jiten.moe.
Stroke order: KanjiVG
Animated stroke-order diagrams use KanjiVG by Ulrich Apel, licensed CC BY-SA 3.0.
Pronunciation audio: VOICEPEAK
Pronunciations are synthesised with VOICEPEAK (AH-Software / Dreamtonics). Shirabe plays only the processed clips hosted on its audio CDN; it never substitutes a device-local voice.
Software
Shirabe is built on Ruby on Rails, Hotwire, and other open-source gems. Its Japanese-language tooling includes:
kabosu: tokenisation
Ruby bindings for Sudachi and its full SudachiDict edition. Shirabe uses it to split and read Japanese text. kabosu on GitHub, Apache-2.0.
daidai: conjugation
Pure-Ruby Japanese verb and adjective conjugation used to derive and explain inflected forms. daidai on GitHub.
Typefaces
Shirabe uses Inter Tight, Newsreader, Noto Sans JP, and JetBrains Mono.
Found something missing? Let us know.