Where Shirabe gets its data
Credits & attribution
Shirabe stands on decades of open-source lexicography. Every entry, reading, and tag here originates from the projects below—please visit them, support them, and check their licences before redistributing.
Source versions
Every dataset we currently have loaded, with its origin and exact upstream build. Re-importing a source updates its row here, so this stays an accurate reference. Licences and fuller attribution follow below.
| Source | Origin | Licence | Version | Data date |
|---|---|---|---|---|
| JMdict — word entries | jmdict-simplified (EDRDG) | EDRDG | 3.6.2 | 2026-06-22 |
| JMnedict — names | jmdict-simplified (EDRDG) | EDRDG | 3.6.2 | 2026-06-22 |
| KANJIDIC2 — kanji | jmdict-simplified (EDRDG) | EDRDG | 3.6.2 | 2026-06-22 |
| KRADFILE — kanji → components | jmdict-simplified (EDRDG) | CC BY-SA 4.0 | — | 2026-06-29 |
| RADKFILE — radical → kanji | jmdict-simplified (EDRDG) | CC BY-SA 4.0 | 3.6.2 | 2026-06-29 |
| Example sentences (Tatoeba) | Tatoeba (sense links: jmdict-simplified) | CC BY 2.0 FR | — | 2026-07-15 |
| Kanken levels (漢検) | mimneko/kanji-data | — | — | 2026-06-26 |
| Jōyō table (常用漢字表 本表) | mimneko/kanji-data | CC0 1.0 | — | 2026-02-10 |
| Jōyō appendix (付表) | mimneko/kanji-data | CC0 1.0 | — | 2026-02-09 |
| Radical names — Kanji alive | kanjialive | CC BY 4.0 | — | 2021-02-24 |
| Pitch accents — UniDic | UniDic (NINJAL) | Modified BSD | 202512 | 2025-12-31 |
| Word frequency — jiten-global | jiten.moe | CC BY-SA 4.0 | — | 2026-07-15 |
| Wikipedia abstracts | DBpedia | CC BY-SA 3.0 | 2016-10 | 2026-06-28 |
| BabelStone IDS — kanji structure | BabelStone (Andrew West) | — | — | 2025-06-27 |
| Stroke order — KanjiVG | KanjiVG (Ulrich Apel) | CC BY-SA 3.0 | — | 2025-08-16 |
Dictionary data
JMdict — Japanese↔multilingual word entries
Compiled by the Electronic Dictionary Research and Development Group (EDRDG, James Breen et al.) and distributed under the EDRDG licence. We use the jmdict-simplified JSON conversion by Stanislav Petrov.
JMnedict — proper-noun dictionary
Also from EDRDG, under the same licence terms, and consumed through the jmdict-simplified JMnedict release.
KANJIDIC2 — kanji dictionary
EDRDG, under the same licence. Shirabe uses jmdict-simplified’s KANJIDIC2 all release, with kanji meanings in English, French, Spanish, and Portuguese.
JLPT levels (N5–N1)
The levels shown on kanji and words come from Jonathan Waller’s JLPT Resources, licensed CC BY 4.0. Because the post-2010 JLPT publishes no official kanji or vocabulary list, these are unofficial lists: kanji from the KANJIDIC snapshot and words through Bluskyo/JLPT_Vocabulary.
KRADFILE — kanji component breakdown
Kanji parts come from EDRDG’s KRADFILE / KRADFILE2 (Michael Raine, James Breen et al.), licensed CC BY-SA 4.0 and consumed through jmdict-simplified.
BabelStone IDS — kanji structure
Positional Ideographic Description Sequences (⿰⿱⿴…) come from Andrew West’s BabelStone IDS data, released to the public domain.
Kanken (漢字検定) levels
Assigned Kanji Kentei levels come from mimneko/kanji-data, released under CC0 1.0.
Jōyō kanji table (常用漢字表)
Official readings and examples come from the Japanese Agency for Cultural Affairs’ 2010 常用漢字表, digitised by mimneko/kanji-data under CC0 1.0. The underlying table is a Japanese cabinet notification and is not subject to copyright.
成り立ち — kanji formation and glyph origin
Formation types and Japanese glyph-origin notes come from Japanese Wiktionary under CC BY-SA 4.0. Shinjitai inherit their kyūjitai’s origin. A few jōyō gaps have an AI-assigned classification and are marked on the kanji page.
RADKFILE — search radicals to kanji
The 253 search radicals and their kanji come from EDRDG’s RADKFILE / RADKFILE2, licensed CC BY-SA 4.0 and consumed through jmdict-simplified.
Kangxi radical names and meanings
Japanese readings, English glosses, stroke counts, and positional categories come from Kanji alive, licensed CC BY 4.0.
Wikipedia abstracts via DBpedia
Lead summaries come from DBpedia’s long_abstracts dataset. The text belongs to its Wikipedia contributors and is dual-licensed under CC BY-SA 3.0 and the GNU Free Documentation License.
Example sentences — Tatoeba
Example sentences come from Tatoeba through the JMdict examples set, licensed CC BY 2.0 FR and consumed through jmdict-simplified.
Pitch accents — UniDic
Pitch-accent data comes from UniDic by the National Institute for Japanese Language and Linguistics, available under GPL 2.0, LGPL 2.1, or Modified BSD.
Word frequency — jiten.moe
Rank ordering for the frequency sort comes from the global frequency list published by jiten.moe.
Stroke order — KanjiVG
Animated stroke-order diagrams use KanjiVG by Ulrich Apel, licensed CC BY-SA 3.0.
Pronunciation audio — AivisSpeech
Pronunciations are synthesised with AivisSpeech using the るな and TANAKA voices under the Aivis Common Model License, which permits commercial use. The browser’s speech synthesis fills in when no clip is available.
Software
Shirabe is built on Ruby on Rails, Hotwire, and other open-source gems. Its Japanese-language tooling includes:
kabosu — tokenisation
Ruby bindings for Sudachi and its full SudachiDict edition. Shirabe uses it to split and read Japanese text. kabosu on GitHub, Apache-2.0.
daidai — conjugation
Pure-Ruby Japanese verb and adjective conjugation used to derive and explain inflected forms. daidai on GitHub.
Typefaces
Shirabe uses Inter Tight, Newsreader, Noto Sans JP, and JetBrains Mono.
Found something missing? Let us know.