Lexical Universe Map

Where does Bitigçi stand in its corpus? Lexical units in the dictionary and those still to come

Where does Bitigçi stand in its corpus?

This map compares the lexical universe of the Bitigçi Corpus (BD) – the words, phrases, idioms and compound verbs actually used in it – with the headwords of the Bitigçi dictionary. Bitigçi sits in the middle as a single blue area; every blue dot is a headword attested in the corpus. Around it, coloured clusters show the candidates not yet in the dictionary, grouped by topic. Region sizes are proportional to counts: as new entries (words, phrases, idioms, compound verbs) are added, the blue area grows and the clusters around it shrink. The growing corpus is re-scanned every 6 hours; new dictionary entries are taken into account every 30 minutes.

Dictionary sync: — · Corpus scan: —

Guide: how to read this page (click to open)

1. Key terms

Lexical unit
A word or fixed word combination that deserves its own dictionary entry: a word (kalem), a derived word, a term or compound noun (kas tonusu), an idiom (göze girmek) or a compound verb (kabul etmek).
Lexical universe
Everything in the corpus that could be a lexical unit, after spelling errors, inflected forms, proper names, abbreviations, foreign words and free combinations have been filtered out.
Candidate
An item in the universe that is not (yet) a Bitigçi headword. Candidates are grouped as single words, derivations, noun phrases, idioms and compound verbs.
Expected lexical units
Not every candidate is a real lexical unit. A random sample from each candidate group was reviewed by lexicographic criteria; the share of genuine units found there is applied to the whole group. This gives the realistic number of entries still missing from Bitigçi.
Coverage
Raw: Bitigçi ÷ (Bitigçi + all candidates). Expected: Bitigçi ÷ (Bitigçi + expected new units) – the realistic measure. Token: share of running-text occurrences that belong to Bitigçi headwords.

2. Reading the map

  • Blue area and dots Bitigçi. Inside it, lighter circles are topic regions (Health, Law …); each dot is a headword, frequent ones in the centre and drawn larger.
  • Coloured clusters around it are candidates outside the dictionary, one cluster per topic, placed next to the matching Bitigçi region. Faint dots show candidate density (each dot ≈ the number given in the legend); filled ringed dots are candidates reviewed and judged real lexical units; hollow rings are the most probable, not yet reviewed units according to the prediction model.
  • Lines: brown = phrase ↔ component word, green = derivation ↔ base, grey curves = links between a candidate cluster and the Bitigçi regions where its components live.
  • Dashed ring inside the blue area: how big Bitigçi was in the first recorded week. The gap between the ring and the edge is the growth since then.
  • Measure: “Expected lexical units” (default) sizes clusters by the realistic estimate; “Raw candidate pool” shows every filtered candidate, including the long tail of noise-prone items.
  • Scope: all, single words only, noun phrases only, or idioms and compound verbs only.
  • Drag to pan, scroll or pinch to zoom; names appear as you zoom in. Click a dot for its links, click a cluster for its figures. The address bar keeps your selection so you can share it.

3. Cards and charts

  • Top cards: dictionary size, corpus size, size of the lexical universe, coverages, candidate pool and the expected number of new units with its range.
  • Coverage by topic: clusters with low coverage are where new entries are most needed; click a row to jump there on the map.
  • Frequency bands: frequent words are almost all in Bitigçi; gaps and noise are concentrated in the long tail.
  • Expected table: pool, reviewed items, share of genuine units with its 95% confidence interval and the expected count for each stratum.
  • Growth: weekly coverage since the first recorded additions, computed against today’s corpus.
  • Estimated list: the expected lexical units themselves. For each group, as many of the most probable candidates as expected, ranked by estimated probability; items without “reviewed” are unverified estimates.

4. Freshness

The corpus keeps growing and new words, phrases, idioms and compound verbs are added to the dictionary every day. The page therefore updates at two speeds: the whole corpus is re-scanned every 6 hours, and new dictionary entries are processed every 30 minutes – a candidate that enters the dictionary moves from the surrounding cluster into the blue area and the coverage and expected figures are recomputed. Both times are shown at the top of the page.

What not to conclude. A small blue area under “Raw candidate pool” does not mean Bitigçi is small: the raw pool contains much noise, the realistic picture is “Expected lexical units”. A candidate is not a decision; every entry still needs definition and evidence work. Coverage is measured against this corpus only; historical, regional and specialised headwords may lie outside it.
Loading the lexical universe…