WSLWordScrambleLab ENไทย

HomeAbout › Methodology & sources

Methodology & sources

Every number, level, definition and pronunciation on this site, and how each one is built, checked and corrected — in one place.

The lexicon in three layers

Results everywhere on this site arrive in the same three layers, in the same order:

LayerWhat it carriesSize
Words with learner dataMeaning, part of speech, syllable break, example, word family and an evidence-graded CEFR level — hand-written, entry by entry.5,028 entries
Other common wordsRecognized and genuinely useful, but no teaching data written yet. Shown plainly and labelled as such.39,562 forms
Rare and technical wordsRecognized for completeness; collapsed by default because they rarely help in a classroom.140,357 forms

The coverage lexicon totals 179,919 recognized 3–12-letter forms. All counts on this page are machine-generated from the dataset at every build and published as JSON at /words/manifest.json?v=6ce1dece. No external dictionary or API is ever queried when you search — recognition runs entirely in your browser.

The honest gap is the first layer: it covers a fraction of what the tool recognizes, and it grows slowly because every entry means writing a definition a learner can actually read. Priority goes to words that appear in the worksheets and words people actually search for.

Where each kind of data comes from

No single dictionary can do all of these jobs, so each kind of information has its own source and its own check.

DataSource and method
Word recognitionThe coverage lexicon described below, built from the English Speller Database (ESDB) and served locally. Described as recognized word forms, not as any game's official dictionary.
CommonnessESDB's own size band, which records how central a word is to the dictionaries ESDB was compiled from. This is dictionary coverage, not corpus frequency — see below.
Definitions and examplesWritten for learners in plain language, then checked against established learner dictionaries (Oxford Learner's Dictionaries, Cambridge Dictionary). Wording is never copied from them.
Part of speech and syllablesAssigned editorially and checked against the same learner dictionaries.
CEFR levelEvidence-graded against three published vocabulary profiles — full detail in the section below.
Pronunciation (IPA)Machine-derived from the CMU Pronouncing Dictionary, General American. Copyright © 1993–2015 Carnegie Mellon University, used under its two-clause BSD licence — full terms and warranty disclaimer.
Pali terms on Buddhist entriesAligned with P.A. Payutto, Dictionary of Buddhism. The English definitions remain this site's own plain-language wording.
Etymology, where usedEtymonline.

The coverage lexicon, exactly

The words this site recognises come from the English Speller Database (ESDB, formerly SCOWL), compiled by Kevin Atkinson from a body of public-domain wordlists. The build is pinned to release rel-2026.02.25 and generated locally from that release, so the same inputs always produce the same list.

Five spellings are generated and merged — American, British -ise, British -ize, Canadian and Australian — because the site is used internationally and no one spelling is treated as the standard. From that union the site keeps only lowercase a–z forms of three to twelve letters: no capitals, apostrophes, hyphens or accents. Four further screens are applied, each using the database's own metadata rather than guesswork about spelling:

  • Abbreviations are removed — forms ESDB records only as abbreviations. A word that is also an ordinary word stays, so art and auto remain.
  • Prefixes and suffixes are removed, because they are not words.
  • Words carried only by the UK Advanced Cryptics Dictionary are removed. That component is almost entirely obscure crossword vocabulary; a word it shares with any other component stays.
  • This site's own exclusion list is applied — 312 offensive forms, kept as an editorial decision of this site rather than a property of whichever wordlist is in use.

The result is 179,919 word forms. Proper nouns never enter it: ESDB records them capitalised, and the lowercase rule excludes them.

The common band is ESDB's size 35 tier — the words carried by even the smallest dictionaries ESDB draws on. This is a measure of dictionary coverage, not of how often a word is used. One consequence is worth stating plainly: a word that every dictionary lists but nobody says any more, such as whilst or hath, counts as common here, and a recent coinage that dictionaries have not caught up with counts as extended. A true frequency measure would need a corpus, and this site does not currently use one.

The screen is deliberate: ethnic and disability slurs are removed outright, and crude sexual and drug slang is kept out of the results shown by default. Ordinary vocabulary that happens to be grim — kill, gun, murder — stays, because it appears in news, literature and exam texts, and removing it would make the tool quietly wrong.

Until 11 September 2026 the recognition layer was a different thing: a community-merged tournament wordlist of Collins Scrabble Words vintage (CSW12-era, on the evidence of the list itself) merged with the public-domain YAWL list. It was replaced because the Collins-derived half carried no licence permitting this site to redistribute it. The change removed about 66,700 obscure forms and added about 6,200; every one of the 5,024 recognised words in the learner layer stayed recognised.

Data licences

The definitions, example sentences, word families, topics and worksheets on this site are its own work. The data underneath them is not, and the licences that permit its use also require these notices.

ComponentLicence and notice
English Speller Database (ESDB), release rel-2026.02.25
the recognition lexicon and its size bands
Copyright © 2000–2026 Kevin Atkinson. Permission to use, copy, modify, distribute and sell any part of the database, or word lists created from it, is granted without fee provided the copyright notice appears in all copies and both that notice and the permission notice appear in supporting documentation. Full notice. ESDB is itself derived from 12dicts and ENABLE2K, both in the public domain; Alan Beale is credited by the project as their author and a principal contributor.
CMU Pronouncing Dictionary
the IPA pronunciations
Copyright © 1993–2015 Carnegie Mellon University, under a two-clause BSD licence. Full terms, conditions and warranty disclaimer.
Typefaces Bricolage Grotesque, JetBrains Mono and Noto Sans Thai under the SIL Open Font Licence 1.1, with their licence texts shipped beside the fonts. The body face is a subset of Adobe's Source Sans 3, also under the OFL; because a subset is a modified version and that licence reserves the name Source, the subset is renamed WSL Sans and Adobe's copyright notice is preserved inside it.

CEFR levels and the five evidence states

CEFR bands describe what a learner can do at a given stage, from A1 for beginners through C2 for near-native users. The Council of Europe framework grades proficiency, not vocabulary — so a "B1 word" is an applied judgement from published vocabulary profiles, not an official classification of the word. Applying bands to individual words is a judgement, and published wordlists disagree with each other at the margins.

Every level in the learner layer is therefore graded by evidence against three published references, matching on word, part of speech and level together:

  • The CEFR-J Wordlist Version 1.5, compiled by Yukio Tono, Tokyo University of Foreign Studies (© Tono Laboratory, TUFS) — A1–B2.
  • The Oxford 3000 and Oxford 5000 by CEFR level (© Oxford University Press) — A1–C1.
  • Octanove Vocabulary Profile C1/C2 (Octanove Labs, CC BY-SA 4.0) — C1–C2.

Each label carries one of five states. Disagreement dominates agreement: if any reference places the word in a different band, the label is disputed even when another reference matches.

MarkStateMeaningEntries
✓✓CorroboratedTwo or more references agree on word, part of speech and level.1,606
ReferencedThe one reference that covers the word agrees; no other covers it.1,115
DisputedThe references place the word in different bands — usually neighbouring ones, because published lists genuinely differ.1,729
✓*ReviewedA conflict examined by hand and decided, with the reasoning recorded. Never assigned by machine.0
DraftCovered by no checked reference yet; drafted from teaching experience.578

The full distribution by band, regenerated at every build:

Band✓✓✓*Total
A155358164093868
A23441254280100997
B129627373101341,434
B233337738401811,275
C18026720059426
C2015201128

Disputed entries sit in a review queue. A review either adopts a reference's level or keeps this site's, and in both cases the reasoning is recorded; only then does a label earn the ✓* mark. Where the wordlists disagree they usually differ by a single band — neighbouring bands genuinely overlap — so a ≠ level should be read as approximate rather than wrong, and any ≠ or ⚠ word is worth checking against a published list before it goes into an official document or a test.

The bands as they stand are published in full at vocabulary by level. Reading a whole band at once is the quickest way to spot a wrong label, and corrections are welcome.

Pronunciation

Transcriptions are machine-derived from the CMU Pronouncing Dictionary and show a General American accent. If you teach a British model, expect differences in words like water and letter, where the American tapped consonant and r-coloring are audible. The transcriptions are broad rather than narrow; stress is marked, but syllable boundaries within a transcription are approximate.

Words with more than one pronunciation are checked by sense: heteronyms such as use, read, tear, lead and live carry the transcription of the sense their entry defines, and a release gate now blocks any build in which one of these regresses. Six wrong-sense transcriptions were found and corrected this way on 23 August 2026.

The speaker button beside each word uses the speech voice built into your own device. No audio is recorded or hosted here. One practical note for Android: pages opened inside the LINE, Facebook or Messenger apps run in a browser view that has no speech engine, so the button explains this and asks you to open the page in Chrome; on iPhone the in-app browser can speak. A synthetic voice is a rough guide rather than a model worth imitating, and no substitute for a teacher's own pronunciation in class.

Change log

A site does not become trustworthy by claiming it makes no mistakes; it becomes trustworthy by showing how mistakes are found and fixed. Dataset changes are recorded here, most recent first.

  • 24 August 2026 — CEFR evidence re-graded from a binary confirmed/draft model into the five states above. Of 965 labels previously counted as confirmed, 385 turned out to have a second reference actively disagreeing and are now openly marked disputed. No levels were changed by the re-grade.
  • 23 August 2026 — six wrong-sense pronunciations corrected (use, read, tear, lead, fly, left) and a heteronym release gate added so they cannot regress. 312 offensive terms removed from the coverage lexicon, matching Collins' own 2021/2022 removals; the lexicon moved from 240,742 to 179,919 recognized forms.
  • 22 August 2026 — first full check of every CEFR label against the published references, under the earlier binary model since superseded.

Report a problem

Every result card in the unscrambler carries a report link for that word. For anything else — a wrong level, a bad definition, a scramble with two valid answers, a worksheet procedure that does not survive contact with a real class — write to wordscramblelab@gmail.com. Corrections from working teachers are worth more to this project than any amount of new content, and confirmed corrections are recorded in the change log above.

More about the people and the thinking behind the site is on the About page; the tools themselves are the word unscrambler and the scramble maker.