const ALIGN_LIMITS
ALIGN_LIMITS: { readonly word: 64; readonly reading: 256; }
What segmentWord will work on: a longer word or reading is refused rather than searched.
@johnmorrisdotca/hikidashi 1.0.0 · 7 entry points · 93 exports
@johnmorrisdotca/hikidashiALIGN_LIMITS buildKanjiTest Deinflection dictionaryForms eraOnDate EraYear eraYearOf eraYearsOf EXTRACT_LIMITS extractFromText ExtractInput ExtractResult followsAsVerbStem formatEraYearJapanese formatEraYearRomaji hiraganaToKatakana JAPANESE_ERAS JapaneseEra kanjiAsWord kanjiCost KanjiDifficultySource kanjiIn KanjiTest KanjiTestQuestion katakanaToHiragana KnownWords LARGEST_JAPANESE_NUMBER LONGEST_ENDING parseEnglishNumber parseEraYear parseJapaneseNumber readJapaneseNumber ReadWord sanitizePastedText SEGMENT_KINDS SegmentKind segmentWord sentenceDifficulty UNKNOWN_KANJI_COST VERSION westernYearFor WORD_CLASSES wordCandidates WordClass wordClassesOf WordSegment writeJapaneseNumber
ALIGN_LIMITS: { readonly word: 64; readonly reading: 256; }
What segmentWord will work on: a longer word or reading is refused rather than searched.
buildKanjiTest(pool: readonly ReadWord[], { seed, count }: { seed: number; count: number; }): KanjiTest
The test: words to write in kanji, then different words to read.
The seed decides which words and in what order, so the test on paper is the test that was on screen and a link sent to somebody is the same test. Each word is asked once, and a word whose reading does not fit it (segmentWord says null) is left out. A list too short for two halves of different words asks its words both ways rather than leave the second half short. A count below one asks nothing.
type Deinflection = { base: string; wordClass: WordClass };
dictionaryForms(written: string): Deinflection[]
Every dictionary form written could be a conjugation of, with the kind of word it would have to be. The written form itself is not among them: a word written as the catalogue lists it is the caller's plain match.
Likeliest first: the reading that accounts for more of the written ending leads, so 食べました is 食べる (ました) before it is 食べます (した). A caller that cannot check the kind of word, like the news reader's lookup, tries them in this order.
eraOnDate(date: string | Date): EraYear | null
The era year a day is in, from YYYY-MM-DD or a Date (read as the day it shows in its own time zone).
Exact where eraYearOf cannot be: 1989-01-07 is 昭和64 and 1989-01-08 is 平成1. Null for a day that is not on the calendar (the 30th of February), or before 明治 began on 1868-10-23.
type EraYear = { era: JapaneseEra; /** The year within the era, counting from one. */ year: number; /** The Western year it falls in. */ westernYear: number; };
An era year that parsed or was found, and the year it names.
eraYearOf(westernYear: number): EraYear | null
The era year a Western year is usually written as: the newest era's, so 1989 is 平成1 and 2019 is 令和1. Null where eraYearsOf has nothing. Use eraYearsOf to get both of a changeover year, or eraOnDate when the day is known.
eraYearsOf(westernYear: number): EraYear[]
Every era year that falls in a Western year, the newest era first.
Most years have one answer. The five years an era changed in have two: 1912 is 大正1 and 明治45, 1926 is 昭和1 and 大正15, 1989 is 平成1 and 昭和64, 2019 is 令和1 and 平成31. 1868 is 明治1 alone. A year before 1868, a year past the reach of the era still running, or a number that is not a whole year gives an empty list.
EXTRACT_LIMITS: { readonly characters: 20000; readonly wordLength: 8; readonly candidates: 4000; }
What a paste is held to.
extractFromText({ text, known }: ExtractInput): ExtractResult
The words and kanji a paste comes to: the words known lists, longest match first, and every kanji in the text.
Both, rather than one or the other. A handout of vocabulary is wanted as words; the same handout is also the kanji a learner has to write, and which of the two somebody meant is not for this to decide. Show the result first, and let them take out what they did not want.
type ExtractInput = { text: string; known?: KnownWords };
What to read, and the dictionary to read it against. Without known only kanji are found.
type ExtractResult = { /** Dictionary forms: 行きます is 行く. */ words: string[]; /** Single kanji, whether or not they belong to a word. */ kanji: string[]; /** What the paste held, for the line that says what happened. */ stats: { characters: number; truncated: boolean; kanji: number; words: number }; };
The words found, then the kanji, each once, in the order they first appear; and what the paste held.
followsAsVerbStem(following: string): boolean
Whether following opens with an ending that makes the word before it a verb stem.
formatEraYearJapanese(found: EraYear, options?: { gannen?: boolean; }): string
The date as it is written in Japanese: 平成3年. With gannen, the first year is written the way forms write it: 令和元年.
formatEraYearRomaji(found: EraYear): string
The date as it is written in Latin letters: Heisei 3.
hiraganaToKatakana(value: string): string
Hiragana written as katakana: ひらがな to ヒラガナ. Anything that is not a hiragana letter is left as it is.
JAPANESE_ERAS: readonly JapaneseEra[]
The five modern eras, oldest first.
明治 is dated from 23 October 1868, when the name was proclaimed; its years are counted as if it began with 1868, so 明治1 is 1868 and 明治45 is 1912. Each later era begins the day the emperor before it died, which is also the day the year number starts again from 元年.
type JapaneseEra = { /** The name in kanji, which is how a date is written. */ kanji: string; /** The Latin name, capitalised as English writes it. */ romaji: string; /** The reading in hiragana, for anyone meeting the kanji for the first time. */ reading: string; /** The Western year the era's first year falls in. */ startYear: number; /** * The highest year the era reached, or null while it is still running. * * 昭…
An era, as it is written, said and counted.
kanjiAsWord(kanji: string, kun: readonly string[], on: readonly string[]): ReadWord | null
A kanji on its own, as a word a test can ask.
A dictionary marks okurigana with a dot - た.べる for 食 - so 食 becomes 食べる read たべる, the question a worksheet would ask. A reading without one is the kanji's own word (牛, うし), and a kanji with no kun reading is asked by its on reading (門, もん). Prefix and suffix forms (-び) are passed over. Null for a kanji with neither reading.
kanjiCost(entry: KanjiDifficultySource | null | undefined): number
The cost of one kanji.
School grades are the scale a family already understands, so grades 1 to 6 cost their own number. Grade 8 is the rest of jōyō, taught in secondary school, and 9 and 10 are the name kanji, which a learner meets late if at all. Ungraded characters fall back to how common they are, since a frequent character is easier than a rare one whatever any curriculum says. A kanji nothing is known about costs UNKNOWN_KANJI_COST, the most.
type KanjiDifficultySource = { /** The school grade: 1 to 6 for elementary, 8 for the rest of the jōyō kanji, 9 and 10 for name kanji (KANJIDIC2's numbering). Null when ungraded. */ grade: number | null; /** How common the kanji is, 1 being the most common. Null when not ranked. */ frequencyRank: number | null; };
What a character costs, given what your dictionary knows about it.
kanjiIn(text: string): string[]
Every kanji in the text, once each, in the order they appear. Kana, punctuation and 々 are not kanji.
type KanjiTest = { writing: KanjiTestQuestion[]; reading: KanjiTestQuestion[] };
A test: words to write in kanji from their readings, then different words to read.
type KanjiTestQuestion = ReadWord & { segments: WordSegment[] };
One question of a test: the word, its reading, and the word cut into its segments.
katakanaToHiragana(value: string): string
Katakana written as hiragana: カタカナ to かたかな. The long-vowel mark ー and anything else that is not a katakana letter are left as they are.
type KnownWords = ReadonlyMap<string, readonly WordClass[]>;
The words your dictionary confirmed, each with the kinds of word it is. A word with no kinds is still a match as written; only a conjugated form needs its dictionary form to be the right kind (wordClassesOf reads the kinds from a dictionary's parts of speech).
LARGEST_JAPANESE_NUMBER: number
The largest number this module will answer for.
JavaScript stops counting exactly above it, and a number the runtime cannot hold is worse than one that is not offered: it would answer, confidently, with a value a digit or two out.
LONGEST_ENDING: number
The longest ending any rule strips, so a caller knows how far past a word to look.
parseEnglishNumber(text: string): number | null
The value of a number written in words, or null when it is not one.
Digits are allowed inside it - "5 man", "20 thousand" - since that is how somebody actually types a magnitude they know the name of. Anything the table does not hold makes the whole thing null rather than a partial reading: "five cats" is not five, and answering as though it were would be a confident wrong answer to a question nobody asked.
parseEraYear(raw: string): EraYear | null
The era year a text names, or null when it does not name one.
元年, "first year", is spelled out rather than numbered on anything official, so it is read as the 1 it means. Everything else is a plain number, half-width or full, with or without the 年 that usually follows it. Latin names may be spelled with a long-vowel macron (Shōwa) or without.
parseJapaneseNumber(text: string): number | null
The value of a number written in Japanese, or null when it is not one.
Mixed spellings are the common case rather than an edge: a headline writes 1億2000万 with digits for the awkward parts and characters for the magnitudes, and a textbook writes the same number 一億二千万. Both are read here, and so is a plain run of digits, because the answer wants to go both ways from whichever one was typed.
The formal numerals are read too (壱万弐千, 萬), and so is a thousands comma. Units must descend: 万億 is not a number anybody wrote on purpose, and reading it as one would put a confident wrong answer on the page. Null for anything else, including a minus sign, a decimal point, a bare 万 and a number past 9,007,199,254,740,991.
readJapaneseNumber(value: number): string | null
The number said aloud, in hiragana, or null when it is out of range.
type ReadWord = { word: string; reading: string };
A word and its reading in hiragana.
sanitizePastedText(raw: string): { text: string; truncated: boolean; }
The text as it will be read: cut to the cap, stripped of what cannot be read.
SEGMENT_KINDS: { readonly kanji: "kanji"; readonly kana: "kana"; readonly okurigana: "okurigana"; }
What a run of a word is: kanji with their part of the reading, kana inside the word, or the kana that end it.
type SegmentKind = (typeof SEGMENT_KINDS)[keyof typeof SEGMENT_KINDS];
segmentWord(word: string, reading: string): WordSegment[] | null
A word cut into its kanji and its kana, each kanji run with its share of the reading.
The kana a word is written with are anchors in its reading: 形が合う read かたちがあう is 形 かたち, が, 合 あ, う. Kana that end a word after its kanji are its okurigana, which the writing half asks for in brackets. A katakana reading, or katakana in the word, reads the same as hiragana. A compound of several kanji (絵日記, えにっき) is one segment, since a reading cannot be split between kanji without a dictionary.
Null when the reading does not fit the word, so that a test can leave the word out rather than mark a wrong answer right. Also null for a word with no kanji, a reading with a space or control character in it, and a word or reading longer than ALIGN_LIMITS.
sentenceDifficulty(text: string, costs: ReadonlyMap<string, number>): number
The sentence's score, lowest first: its length in characters plus three times its hardest kanji.
Kana-only sentences score their length alone, which puts them at the top where they belong: a beginner can read every one of them. A kanji missing from costs counts as the hardest kind.
UNKNOWN_KANJI_COST: 20
A character nobody has graded and nobody counts as frequent.
VERSION: "1.0.0"
The package's version.
westernYearFor(era: JapaneseEra, year: number): number | null
The Western year an era year falls in, or null when the era never had that year.
Minus one, because an era's first year is year one rather than year zero: 平成1年 is 1989, the year the era began, not the year after it. 昭和65 is null, since 昭和 ended in its 64th year.
WORD_CLASSES: { readonly godan: "godan"; readonly ichidan: "ichidan"; readonly iAdjective: "iAdjective"; readonly suru: "suru"; readonly kuru: "kuru"; }
活用: the dictionary form a conjugated word was written from.
A sentence writes its verbs and adjectives conjugated: 行きます, 食べました, 待って, 高くて. A dictionary lists them as 行く, 食べる, 待つ and 高い. Look a sentence up as it is written and 先生と学校へ行きます finds 行き, a noun meaning "bound for" that is also the stem of the verb, instead of 行く.
This walks a written form back through its endings, one at a time, the way a learner does: 行きませんでした is the polite negative past of 行く; 食べなかった is the past of 食べない, which is the negative of 食べる. Every answer says which kind of word it would have to be, and a caller keeps only the ones its dictionary confirms are that kind, so 行き is never taken for a verb stem when nothing follows it and a noun ending in る is never taken for a verb.
There is deliberately no rule that reads a bare stem as its verb. 東京行き is a noun, and 行き alone is only ever a verb when an ending follows it, which is exactly when a longer rule matches.
Covered: godan and ichidan verbs, い adjectives, and the two irregular verbs する and 来る, in the polite, negative, past, te, conditional, volitional, passive, causative, potential and "wanting" forms and chains of them. Not covered: the imperative, keigo verbs, and な adjectives (which do not conjugate onto themselves).
wordCandidates(text: string): string[]
Every substring the dictionary might know, so one question can be asked about the whole paste: look all of them up at once, and pass the ones found as known. Capped at EXTRACT_LIMITS.candidates, because a long paste of Japanese has more substrings than anybody needs to look up. Conjugated forms are offered as the dictionary forms they could come from, and a single character is never offered, since that is a kanji.
type WordClass = (typeof WORD_CLASSES)[keyof typeof WORD_CLASSES];
wordClassesOf(characters: string, partsOfSpeech: readonly string[]): WordClass[]
The kinds of word a dictionary entry is, from the parts of speech it lists, so a conjugated form is only taken for a verb or an adjective the dictionary says is one.
Reads plain labels ("godan verb", "ichidan verb", "i-adjective" or "い adjective", "suru verb" or "する verb"), and JMdict's tags (v5k, v1, adj-i, vs-i, vk), in any case. A dictionary that lists 来る only as "an intransitive verb" is covered: a verb whose spelling ends in 来る or くる is a 来る verb.
type WordSegment = { text: string; kind: SegmentKind; reading: string };
One run of a word, and its share of the reading in hiragana.
writeJapaneseNumber(value: number, options?: { formal?: boolean; }): string | null
The number as Japanese writes it, or null when it is out of range (negative, not a whole number, or past 9,007,199,254,740,991).
Grouped by myriads, so each group is a number below 10,000 followed by its unit: 120,000,000 comes out 一億二千万 rather than as a run of digits. With formal, the figures that can be altered are written in their formal characters, as on a contract or a banknote: 壱万 for 一万, 弐千 for 二千.
@johnmorrisdotca/hikidashi/warekieraOnDate EraYear eraYearOf eraYearsOf formatEraYearJapanese formatEraYearRomaji JAPANESE_ERAS JapaneseEra parseEraYear westernYearFor
eraOnDate(date: string | Date): EraYear | null
The era year a day is in, from YYYY-MM-DD or a Date (read as the day it shows in its own time zone).
Exact where eraYearOf cannot be: 1989-01-07 is 昭和64 and 1989-01-08 is 平成1. Null for a day that is not on the calendar (the 30th of February), or before 明治 began on 1868-10-23.
type EraYear = { era: JapaneseEra; /** The year within the era, counting from one. */ year: number; /** The Western year it falls in. */ westernYear: number; };
An era year that parsed or was found, and the year it names.
eraYearOf(westernYear: number): EraYear | null
The era year a Western year is usually written as: the newest era's, so 1989 is 平成1 and 2019 is 令和1. Null where eraYearsOf has nothing. Use eraYearsOf to get both of a changeover year, or eraOnDate when the day is known.
eraYearsOf(westernYear: number): EraYear[]
Every era year that falls in a Western year, the newest era first.
Most years have one answer. The five years an era changed in have two: 1912 is 大正1 and 明治45, 1926 is 昭和1 and 大正15, 1989 is 平成1 and 昭和64, 2019 is 令和1 and 平成31. 1868 is 明治1 alone. A year before 1868, a year past the reach of the era still running, or a number that is not a whole year gives an empty list.
formatEraYearJapanese(found: EraYear, options?: { gannen?: boolean; }): string
The date as it is written in Japanese: 平成3年. With gannen, the first year is written the way forms write it: 令和元年.
formatEraYearRomaji(found: EraYear): string
The date as it is written in Latin letters: Heisei 3.
JAPANESE_ERAS: readonly JapaneseEra[]
The five modern eras, oldest first.
明治 is dated from 23 October 1868, when the name was proclaimed; its years are counted as if it began with 1868, so 明治1 is 1868 and 明治45 is 1912. Each later era begins the day the emperor before it died, which is also the day the year number starts again from 元年.
type JapaneseEra = { /** The name in kanji, which is how a date is written. */ kanji: string; /** The Latin name, capitalised as English writes it. */ romaji: string; /** The reading in hiragana, for anyone meeting the kanji for the first time. */ reading: string; /** The Western year the era's first year falls in. */ startYear: number; /** * The highest year the era reached, or null while it is still running. * * 昭…
An era, as it is written, said and counted.
parseEraYear(raw: string): EraYear | null
The era year a text names, or null when it does not name one.
元年, "first year", is spelled out rather than numbered on anything official, so it is read as the 1 it means. Everything else is a plain number, half-width or full, with or without the 年 that usually follows it. Latin names may be spelled with a long-vowel macron (Shōwa) or without.
westernYearFor(era: JapaneseEra, year: number): number | null
The Western year an era year falls in, or null when the era never had that year.
Minus one, because an era's first year is year one rather than year zero: 平成1年 is 1989, the year the era began, not the year after it. 昭和65 is null, since 昭和 ended in its 64th year.
@johnmorrisdotca/hikidashi/numeralsLARGEST_JAPANESE_NUMBER parseEnglishNumber parseJapaneseNumber readJapaneseNumber writeJapaneseNumber
LARGEST_JAPANESE_NUMBER: number
The largest number this module will answer for.
JavaScript stops counting exactly above it, and a number the runtime cannot hold is worse than one that is not offered: it would answer, confidently, with a value a digit or two out.
parseEnglishNumber(text: string): number | null
The value of a number written in words, or null when it is not one.
Digits are allowed inside it - "5 man", "20 thousand" - since that is how somebody actually types a magnitude they know the name of. Anything the table does not hold makes the whole thing null rather than a partial reading: "five cats" is not five, and answering as though it were would be a confident wrong answer to a question nobody asked.
parseJapaneseNumber(text: string): number | null
The value of a number written in Japanese, or null when it is not one.
Mixed spellings are the common case rather than an edge: a headline writes 1億2000万 with digits for the awkward parts and characters for the magnitudes, and a textbook writes the same number 一億二千万. Both are read here, and so is a plain run of digits, because the answer wants to go both ways from whichever one was typed.
The formal numerals are read too (壱万弐千, 萬), and so is a thousands comma. Units must descend: 万億 is not a number anybody wrote on purpose, and reading it as one would put a confident wrong answer on the page. Null for anything else, including a minus sign, a decimal point, a bare 万 and a number past 9,007,199,254,740,991.
readJapaneseNumber(value: number): string | null
The number said aloud, in hiragana, or null when it is out of range.
writeJapaneseNumber(value: number, options?: { formal?: boolean; }): string | null
The number as Japanese writes it, or null when it is out of range (negative, not a whole number, or past 9,007,199,254,740,991).
Grouped by myriads, so each group is a number below 10,000 followed by its unit: 120,000,000 comes out 一億二千万 rather than as a run of digits. With formal, the figures that can be altered are written in their formal characters, as on a contract or a banknote: 壱万 for 一万, 弐千 for 二千.
@johnmorrisdotca/hikidashi/deinflectDeinflection dictionaryForms followsAsVerbStem LONGEST_ENDING WORD_CLASSES WordClass wordClassesOf
type Deinflection = { base: string; wordClass: WordClass };
dictionaryForms(written: string): Deinflection[]
Every dictionary form written could be a conjugation of, with the kind of word it would have to be. The written form itself is not among them: a word written as the catalogue lists it is the caller's plain match.
Likeliest first: the reading that accounts for more of the written ending leads, so 食べました is 食べる (ました) before it is 食べます (した). A caller that cannot check the kind of word, like the news reader's lookup, tries them in this order.
followsAsVerbStem(following: string): boolean
Whether following opens with an ending that makes the word before it a verb stem.
LONGEST_ENDING: number
The longest ending any rule strips, so a caller knows how far past a word to look.
WORD_CLASSES: { readonly godan: "godan"; readonly ichidan: "ichidan"; readonly iAdjective: "iAdjective"; readonly suru: "suru"; readonly kuru: "kuru"; }
活用: the dictionary form a conjugated word was written from.
A sentence writes its verbs and adjectives conjugated: 行きます, 食べました, 待って, 高くて. A dictionary lists them as 行く, 食べる, 待つ and 高い. Look a sentence up as it is written and 先生と学校へ行きます finds 行き, a noun meaning "bound for" that is also the stem of the verb, instead of 行く.
This walks a written form back through its endings, one at a time, the way a learner does: 行きませんでした is the polite negative past of 行く; 食べなかった is the past of 食べない, which is the negative of 食べる. Every answer says which kind of word it would have to be, and a caller keeps only the ones its dictionary confirms are that kind, so 行き is never taken for a verb stem when nothing follows it and a noun ending in る is never taken for a verb.
There is deliberately no rule that reads a bare stem as its verb. 東京行き is a noun, and 行き alone is only ever a verb when an ending follows it, which is exactly when a longer rule matches.
Covered: godan and ichidan verbs, い adjectives, and the two irregular verbs する and 来る, in the polite, negative, past, te, conditional, volitional, passive, causative, potential and "wanting" forms and chains of them. Not covered: the imperative, keigo verbs, and な adjectives (which do not conjugate onto themselves).
type WordClass = (typeof WORD_CLASSES)[keyof typeof WORD_CLASSES];
wordClassesOf(characters: string, partsOfSpeech: readonly string[]): WordClass[]
The kinds of word a dictionary entry is, from the parts of speech it lists, so a conjugated form is only taken for a verb or an adjective the dictionary says is one.
Reads plain labels ("godan verb", "ichidan verb", "i-adjective" or "い adjective", "suru verb" or "する verb"), and JMdict's tags (v5k, v1, adj-i, vs-i, vk), in any case. A dictionary that lists 来る only as "an intransitive verb" is covered: a verb whose spelling ends in 来る or くる is a 来る verb.
@johnmorrisdotca/hikidashi/alignALIGN_LIMITS buildKanjiTest hiraganaToKatakana kanjiAsWord KanjiTest KanjiTestQuestion katakanaToHiragana ReadWord SEGMENT_KINDS SegmentKind segmentWord WordSegment
ALIGN_LIMITS: { readonly word: 64; readonly reading: 256; }
What segmentWord will work on: a longer word or reading is refused rather than searched.
buildKanjiTest(pool: readonly ReadWord[], { seed, count }: { seed: number; count: number; }): KanjiTest
The test: words to write in kanji, then different words to read.
The seed decides which words and in what order, so the test on paper is the test that was on screen and a link sent to somebody is the same test. Each word is asked once, and a word whose reading does not fit it (segmentWord says null) is left out. A list too short for two halves of different words asks its words both ways rather than leave the second half short. A count below one asks nothing.
hiraganaToKatakana(value: string): string
Hiragana written as katakana: ひらがな to ヒラガナ. Anything that is not a hiragana letter is left as it is.
kanjiAsWord(kanji: string, kun: readonly string[], on: readonly string[]): ReadWord | null
A kanji on its own, as a word a test can ask.
A dictionary marks okurigana with a dot - た.べる for 食 - so 食 becomes 食べる read たべる, the question a worksheet would ask. A reading without one is the kanji's own word (牛, うし), and a kanji with no kun reading is asked by its on reading (門, もん). Prefix and suffix forms (-び) are passed over. Null for a kanji with neither reading.
type KanjiTest = { writing: KanjiTestQuestion[]; reading: KanjiTestQuestion[] };
A test: words to write in kanji from their readings, then different words to read.
type KanjiTestQuestion = ReadWord & { segments: WordSegment[] };
One question of a test: the word, its reading, and the word cut into its segments.
katakanaToHiragana(value: string): string
Katakana written as hiragana: カタカナ to かたかな. The long-vowel mark ー and anything else that is not a katakana letter are left as they are.
type ReadWord = { word: string; reading: string };
A word and its reading in hiragana.
SEGMENT_KINDS: { readonly kanji: "kanji"; readonly kana: "kana"; readonly okurigana: "okurigana"; }
What a run of a word is: kanji with their part of the reading, kana inside the word, or the kana that end it.
type SegmentKind = (typeof SEGMENT_KINDS)[keyof typeof SEGMENT_KINDS];
segmentWord(word: string, reading: string): WordSegment[] | null
A word cut into its kanji and its kana, each kanji run with its share of the reading.
The kana a word is written with are anchors in its reading: 形が合う read かたちがあう is 形 かたち, が, 合 あ, う. Kana that end a word after its kanji are its okurigana, which the writing half asks for in brackets. A katakana reading, or katakana in the word, reads the same as hiragana. A compound of several kanji (絵日記, えにっき) is one segment, since a reading cannot be split between kanji without a dictionary.
Null when the reading does not fit the word, so that a test can leave the word out rather than mark a wrong answer right. Also null for a word with no kanji, a reading with a space or control character in it, and a word or reading longer than ALIGN_LIMITS.
type WordSegment = { text: string; kind: SegmentKind; reading: string };
One run of a word, and its share of the reading in hiragana.
@johnmorrisdotca/hikidashi/extractEXTRACT_LIMITS extractFromText ExtractInput ExtractResult KnownWords sanitizePastedText wordCandidates
EXTRACT_LIMITS: { readonly characters: 20000; readonly wordLength: 8; readonly candidates: 4000; }
What a paste is held to.
extractFromText({ text, known }: ExtractInput): ExtractResult
The words and kanji a paste comes to: the words known lists, longest match first, and every kanji in the text.
Both, rather than one or the other. A handout of vocabulary is wanted as words; the same handout is also the kanji a learner has to write, and which of the two somebody meant is not for this to decide. Show the result first, and let them take out what they did not want.
type ExtractInput = { text: string; known?: KnownWords };
What to read, and the dictionary to read it against. Without known only kanji are found.
type ExtractResult = { /** Dictionary forms: 行きます is 行く. */ words: string[]; /** Single kanji, whether or not they belong to a word. */ kanji: string[]; /** What the paste held, for the line that says what happened. */ stats: { characters: number; truncated: boolean; kanji: number; words: number }; };
The words found, then the kanji, each once, in the order they first appear; and what the paste held.
type KnownWords = ReadonlyMap<string, readonly WordClass[]>;
The words your dictionary confirmed, each with the kinds of word it is. A word with no kinds is still a match as written; only a conjugated form needs its dictionary form to be the right kind (wordClassesOf reads the kinds from a dictionary's parts of speech).
sanitizePastedText(raw: string): { text: string; truncated: boolean; }
The text as it will be read: cut to the cap, stripped of what cannot be read.
wordCandidates(text: string): string[]
Every substring the dictionary might know, so one question can be asked about the whole paste: look all of them up at once, and pass the ones found as known. Capped at EXTRACT_LIMITS.candidates, because a long paste of Japanese has more substrings than anybody needs to look up. Conjugated forms are offered as the dictionary forms they could come from, and a single character is never offered, since that is a kanji.
@johnmorrisdotca/hikidashi/difficultykanjiCost KanjiDifficultySource kanjiIn sentenceDifficulty UNKNOWN_KANJI_COST
kanjiCost(entry: KanjiDifficultySource | null | undefined): number
The cost of one kanji.
School grades are the scale a family already understands, so grades 1 to 6 cost their own number. Grade 8 is the rest of jōyō, taught in secondary school, and 9 and 10 are the name kanji, which a learner meets late if at all. Ungraded characters fall back to how common they are, since a frequent character is easier than a rare one whatever any curriculum says. A kanji nothing is known about costs UNKNOWN_KANJI_COST, the most.
type KanjiDifficultySource = { /** The school grade: 1 to 6 for elementary, 8 for the rest of the jōyō kanji, 9 and 10 for name kanji (KANJIDIC2's numbering). Null when ungraded. */ grade: number | null; /** How common the kanji is, 1 being the most common. Null when not ranked. */ frequencyRank: number | null; };
What a character costs, given what your dictionary knows about it.
kanjiIn(text: string): string[]
Every kanji in the text, once each, in the order they appear. Kana, punctuation and 々 are not kanji.
sentenceDifficulty(text: string, costs: ReadonlyMap<string, number>): number
The sentence's score, lowest first: its length in characters plus three times its hardest kanji.
Kana-only sentences score their length alone, which puts them at the top where they belong: a beginner can read every one of them. A kanji missing from costs counts as the hardest kind.
UNKNOWN_KANJI_COST: 20
A character nobody has graded and nobody counts as frequent.
この日本語は、まだ日本語を母語とする方の確認を受けていません。訂正を歓迎します。