shirabe.org

Where Shirabe gets its data

Credits & attribution

Shirabe stands on decades of open-source lexicography. Every entry, reading, and tag here originates from the projects below—please visit them, support them, and check their licences before redistributing.

Source versions

Every dataset we currently have loaded, with its origin and exact upstream build. Re-importing a source updates its row here, so this stays an accurate reference. Licences and fuller attribution follow below.

Source Origin Licence Version Data date
JMdict — word entries jmdict-simplified (EDRDG) EDRDG 3.6.2 2026-06-22
JMnedict — names jmdict-simplified (EDRDG) EDRDG 3.6.2 2026-06-22
KANJIDIC2 — kanji jmdict-simplified (EDRDG) EDRDG 3.6.2 2026-06-22
KRADFILE — kanji → components jmdict-simplified (EDRDG) CC BY-SA 4.0 2026-06-29
RADKFILE — radical → kanji jmdict-simplified (EDRDG) CC BY-SA 4.0 3.6.2 2026-06-29
Example sentences (Tatoeba) Tatoeba (sense links: jmdict-simplified) CC BY 2.0 FR 2026-07-15
Kanken levels (漢検) mimneko/kanji-data 2026-06-26
Jōyō table (常用漢字表 本表) mimneko/kanji-data CC0 1.0 2026-02-10
Jōyō appendix (付表) mimneko/kanji-data CC0 1.0 2026-02-09
Radical names — Kanji alive kanjialive CC BY 4.0 2021-02-24
Pitch accents — UniDic UniDic (NINJAL) Modified BSD 202512 2025-12-31
Word frequency — jiten-global jiten.moe CC BY-SA 4.0 2026-07-15
Wikipedia abstracts DBpedia CC BY-SA 3.0 2016-10 2026-06-28
BabelStone IDS — kanji structure BabelStone (Andrew West) 2025-06-27
Stroke order — KanjiVG KanjiVG (Ulrich Apel) CC BY-SA 3.0 2025-08-16

Dictionary data

JMdict — Japanese↔multilingual word entries

Compiled by the Electronic Dictionary Research and Development Group (EDRDG, James Breen et al.) and distributed under the EDRDG licence. We use the jmdict-simplified JSON conversion by Stanislav Petrov.

JMnedict — proper-noun dictionary

Also from EDRDG, under the same licence terms, and consumed through the jmdict-simplified JMnedict release.

KANJIDIC2 — kanji dictionary

EDRDG, under the same licence. Shirabe uses jmdict-simplified’s KANJIDIC2 all release, with kanji meanings in English, French, Spanish, and Portuguese.

JLPT levels (N5–N1)

The levels shown on kanji and words come from Jonathan Waller’s JLPT Resources, licensed CC BY 4.0. Because the post-2010 JLPT publishes no official kanji or vocabulary list, these are unofficial lists: kanji from the KANJIDIC snapshot and words through Bluskyo/JLPT_Vocabulary.

KRADFILE — kanji component breakdown

Kanji parts come from EDRDG’s KRADFILE / KRADFILE2 (Michael Raine, James Breen et al.), licensed CC BY-SA 4.0 and consumed through jmdict-simplified.

BabelStone IDS — kanji structure

Positional Ideographic Description Sequences (⿰⿱⿴…) come from Andrew West’s BabelStone IDS data, released to the public domain.

Kanken (漢字検定) levels

Assigned Kanji Kentei levels come from mimneko/kanji-data, released under CC0 1.0.

Jōyō kanji table (常用漢字表)

Official readings and examples come from the Japanese Agency for Cultural Affairs’ 2010 常用漢字表, digitised by mimneko/kanji-data under CC0 1.0. The underlying table is a Japanese cabinet notification and is not subject to copyright.

成り立ち — kanji formation and glyph origin

Formation types and Japanese glyph-origin notes come from Japanese Wiktionary under CC BY-SA 4.0. Shinjitai inherit their kyūjitai’s origin. A few jōyō gaps have an AI-assigned classification and are marked on the kanji page.

RADKFILE — search radicals to kanji

The 253 search radicals and their kanji come from EDRDG’s RADKFILE / RADKFILE2, licensed CC BY-SA 4.0 and consumed through jmdict-simplified.

Kangxi radical names and meanings

Japanese readings, English glosses, stroke counts, and positional categories come from Kanji alive, licensed CC BY 4.0.

Wikipedia abstracts via DBpedia

Lead summaries come from DBpedia’s long_abstracts dataset. The text belongs to its Wikipedia contributors and is dual-licensed under CC BY-SA 3.0 and the GNU Free Documentation License.

Example sentences — Tatoeba

Example sentences come from Tatoeba through the JMdict examples set, licensed CC BY 2.0 FR and consumed through jmdict-simplified.

Pitch accents — UniDic

Pitch-accent data comes from UniDic by the National Institute for Japanese Language and Linguistics, available under GPL 2.0, LGPL 2.1, or Modified BSD.

Word frequency — jiten.moe

Rank ordering for the frequency sort comes from the global frequency list published by jiten.moe.

Stroke order — KanjiVG

Animated stroke-order diagrams use KanjiVG by Ulrich Apel, licensed CC BY-SA 3.0.

Pronunciation audio — AivisSpeech

Pronunciations are synthesised with AivisSpeech using the るな and TANAKA voices under the Aivis Common Model License, which permits commercial use. The browser’s speech synthesis fills in when no clip is available.

Software

Shirabe is built on Ruby on Rails, Hotwire, and other open-source gems. Its Japanese-language tooling includes:

kabosu — tokenisation

Ruby bindings for Sudachi and its full SudachiDict edition. Shirabe uses it to split and read Japanese text. kabosu on GitHub, Apache-2.0.

daidai — conjugation

Pure-Ruby Japanese verb and adjective conjugation used to derive and explain inflected forms. daidai on GitHub.

Typefaces

Shirabe uses Inter Tight, Newsreader, Noto Sans JP, and JetBrains Mono.

Found something missing? Let us know.

Legend

What the coloured tags mean