Where does Bitigçi stand in its corpus?
This map compares the lexical universe of the Bitigçi Corpus (BD) – the words, phrases, idioms and compound verbs actually used in it – with the headwords of the Bitigçi dictionary. Bitigçi sits in the middle as a single blue area; every blue dot is a headword attested in the corpus. Around it, coloured clusters show the candidates not yet in the dictionary, grouped by topic. Region sizes are proportional to counts: as new entries (words, phrases, idioms, compound verbs) are added, the blue area grows and the clusters around it shrink. The growing corpus is re-scanned every 6 hours; new dictionary entries are taken into account every 30 minutes.
Dictionary sync: — · Corpus scan: —
Guide: how to read this page (click to open)
1. Key terms
- Lexical unit
- A word or fixed word combination that deserves its own dictionary entry: a word (kalem), a derived word, a term or compound noun (kas tonusu), an idiom (göze girmek) or a compound verb (kabul etmek).
- Lexical universe
- Everything in the corpus that could be a lexical unit, after spelling errors, inflected forms, proper names, abbreviations, foreign words and free combinations have been filtered out.
- Candidate
- An item in the universe that is not (yet) a Bitigçi headword. Candidates are grouped as single words, derivations, noun phrases, idioms and compound verbs.
- Expected lexical units
- Not every candidate is a real lexical unit. A random sample from each candidate group was reviewed by lexicographic criteria; the share of genuine units found there is applied to the whole group. This gives the realistic number of entries still missing from Bitigçi.
- Coverage
- Raw: Bitigçi ÷ (Bitigçi + all candidates). Expected: Bitigçi ÷ (Bitigçi + expected new units) – the realistic measure. Token: share of running-text occurrences that belong to Bitigçi headwords.
2. Reading the map
- Blue area and dots Bitigçi. Inside it, lighter circles are topic regions (Health, Law …); each dot is a headword, frequent ones in the centre and drawn larger.
- Coloured clusters around it are candidates outside the dictionary, one cluster per topic, placed next to the matching Bitigçi region. Faint dots show candidate density (each dot ≈ the number given in the legend); filled ringed dots are candidates reviewed and judged real lexical units; hollow rings are the most probable, not yet reviewed units according to the prediction model.
- Lines: brown = phrase ↔ component word, green = derivation ↔ base, grey curves = links between a candidate cluster and the Bitigçi regions where its components live.
- Dashed ring inside the blue area: how big Bitigçi was in the first recorded week. The gap between the ring and the edge is the growth since then.
- Measure: “Expected lexical units” (default) sizes clusters by the realistic estimate; “Raw candidate pool” shows every filtered candidate, including the long tail of noise-prone items.
- Scope: all, single words only, noun phrases only, or idioms and compound verbs only.
- Drag to pan, scroll or pinch to zoom; names appear as you zoom in. Click a dot for its links, click a cluster for its figures. The address bar keeps your selection so you can share it.
3. Cards and charts
- Top cards: dictionary size, corpus size, size of the lexical universe, coverages, candidate pool and the expected number of new units with its range.
- Coverage by topic: clusters with low coverage are where new entries are most needed; click a row to jump there on the map.
- Frequency bands: frequent words are almost all in Bitigçi; gaps and noise are concentrated in the long tail.
- Expected table: pool, reviewed items, share of genuine units with its 95% confidence interval and the expected count for each stratum.
- Growth: weekly coverage since the first recorded additions, computed against today’s corpus.
- Estimated list: the expected lexical units themselves. For each group, as many of the most probable candidates as expected, ranked by estimated probability; items without “reviewed” are unverified estimates.
4. Freshness
The corpus keeps growing and new words, phrases, idioms and compound verbs are added to the dictionary every day. The page therefore updates at two speeds: the whole corpus is re-scanned every 6 hours, and new dictionary entries are processed every 30 minutes – a candidate that enters the dictionary moves from the surrounding cluster into the blue area and the coverage and expected figures are recomputed. Both times are shown at the top of the page.
Lexical universe map
Drag to pan, scroll or pinch to zoom; names appear as you zoom in. The map shows a topic-proportional sample of Bitigçi headwords attested in the corpus and the reviewed candidates; region sizes reflect the full counts.
Bitigçi coverage by topic
Share of Bitigçi headwords in each topic cluster under the measure selected on the map. Clusters with lower coverage are priority areas for new entries. Measure:
From corpus to dictionary: nested layers
Every single form that occurs at least ten times in the corpus is first filtered: spelling errors, forms that lost Turkish characters, inflected forms left by lemmatisation, proper names, abbreviations, foreign words, OCR noise and items already rejected by lexicographers fall outside the universe. Squares are proportional to their areas.
How many more lexical units could enter Bitigçi?
Candidate pools are split into strata by group (single word, derivation, noun phrase, idiom, compound verb), frequency band and cohesion or length. A random sample from each stratum was examined with lexicographic criteria; the share of genuine lexical units found in the sample is applied to the whole stratum. The central estimate is the sum over strata; the range comes from the 95% confidence intervals.
Lexical units expected to enter: estimated list
The candidates most likely to be genuine lexical units, one by one, for each group. A prediction model learned from the reviewed items (their frequency, pattern, topic evidence, cohesion and form) gives every candidate a probability; for each group the list contains as many items as the expected number above, most probable first. Items marked “reviewed” were judged lexical units in the audit; all others are estimates that have not been reviewed yet and may contain noise. Click an item to see its corpus evidence.
Corpus and Bitigçi by frequency band
Almost every very frequent word is already in Bitigçi; coverage drops as frequency falls. The long tail holds both genuine lexical units missing from the dictionary and most of the noise. The chart shows the share of each class within a band; the table gives raw counts.
How is Bitigçi’s coverage growing?
Lines are computed retrospectively against today’s corpus: at the end of each week, how much of today’s lexical universe did the dictionary then cover? Bars show headwords added that week.
Prominent candidates in the corpus
Frequent items outside the dictionary that were judged genuine lexical units in the latest review, filtered for public display. Being listed does not mean an item will enter the dictionary.
Bitigçi headwords in the corpus
The dictionary also reaches beyond the corpus: historical, regional, rare or specialised headwords may be unattested in BD. “Base” is the initial vocabulary; “added” are entries created by Bitigçi’s own work.
Cohesion thresholds for phrases
For noun phrases the tendency of two words to occur together is measured with logDice: 14 + log₂(2·f(xy) / (f(x) + f(y))); — of Bitigçi’s two-word noun phrases found in the corpus lie above the red line. Idioms and compound verbs are built with very frequent verbs (etmek, olmak, gelmek), so logDice undervalues them; for them the directional cohesion is used: the share of the rarer component’s occurrences that fall inside the phrase.
What stays outside the lexical universe
Not every form seen in the corpus is a lexical unit. These classes are counted but kept out of the universe; the figures show the scale of this filtering.
Method, definitions and limitations
Data
Lemma frequencies of the Bitigçi Corpus, two- and three-word lemma sequence counts, the Bitigçi headword list, the new-word queue and the lexicographers’ rejection memory, all read-only from the live corpus database.
Single words
Letter-bearing lemmas with at least 10 occurrences. Headwords (including spelling variants in parentheses) form the Bitigçi layer. The rest are classified: inflected forms (including verb conjugations and forms with consonant alternation or vowel loss), spelling errors, merged spellings, proper names and abbreviations (capitalisation profile in a sentence sample), foreign words and OCR noise, items rejected earlier. The remainder are candidates; words formed by a dictionary stem plus a derivational suffix are shown separately as derivation candidates.
Noun phrases
Two-word lemma sequences with at least 20 occurrences, not ending in a verb, with Turkish parts; sequences with function words, proper names or foreign words are excluded; logDice ≥ 2.
Idioms and compound verbs
Two- and three-word lemma sequences ending in a verb with at least 20 occurrences. Those ending in an auxiliary verb (etmek, olmak, eylemek, kılmak, buyurmak and their passive/causative forms) are compound verbs, the rest idioms or idiom-like verb phrases. They are admitted with the directional-cohesion thresholds above, set so that about 80% of Bitigçi’s own idioms and compound verbs pass. Three-word sequences that merely extend a dictionary phrase are left out.
Topic clusters
The clusters are the corpus-learned domains of the Corpus Statistics page. Sentences in a 5% block sample are labelled with domain anchor words; a word is assigned to the domain where it is concentrated, frequent words spread over many domains form General language, sparse words are distributed by the profile of their frequency band (an estimate). A phrase takes the topic of its rarer, more distinctive component.
Limitations
- Lemmatisation errors can make some headwords look unattested and leave some inflected forms among candidates.
- Phrases of four or more words are not measured; the phrase estimates are therefore lower bounds.
- Sample-based topic labels and estimates carry sampling error; figures are recomputed nightly and the audit is repeated periodically.
Sample: — of the corpus (— sentences), — anchor words, direct topic evidence for — of universe items, noun-phrase threshold logDice ≥ —, — forms set aside as spelling errors. Audit: —.