How the AI mock examiner works
This is an AI-generated practice estimate, not an ICAO Language Proficiency assessment. It cannot award, predict, or guarantee any ICAO level or any test result. Only a licensed rater at an approved test centre can assess your ICAO level.
This page is the whole method. What the six criteria are, how a level is produced, what we have measured about our own marking and published, what is not measured at all, and the case where we give you no level and hand the session back. If you read one section, read the one about what we do when the machine cannot hear you properly. It is the part of this that is new.
What it is, and what it is not
It is a practice tool. It gives you a realistic interview shape, a clock, questions you have not rehearsed, and a report that quotes your own words back at you so you can see where your language stopped working. That is the useful part, and it is the part no amount of reading a syllabus gives you.
It is not an assessment. It has no official standing, it is not affiliated with any test provider or licensing authority, and it cannot award, predict or guarantee a level or a test result. We do not publish a pass rate, because we would have no way of knowing one and because a practice tool that claimed one would be selling you a number it made up. Only a licensed rater at an approved test centre can assess your ICAO level.
The six criteria, and how each one is judged
Every session is judged on the same six criteria. Each one in your report has to arrive with one specific thing to practise, and the examiner is required to support it with two or three short quotations of what you actually said. A score with no quotation behind it is not feedback, because you cannot act on it, and the figures further down say how often that requirement was actually met when we measured it.
Pronunciation
Pronunciation, stress, rhythm and intonation, and how far a first language accent interferes with an international listener understanding you. Having an accent is not the thing being measured. Being hard to understand is.
From a transcript: This is the criterion a transcript supports least well, and the report says so on every session rather than hiding it. The examiner cannot hear you. It can only use what survives into text: words that came through garbled, digits that arrived wrong, places where you corrected yourself, and points where the examiner had to rephrase. The absence of those is not evidence of good pronunciation. It is the absence of evidence, so where there is nothing to report this criterion is marked as not judged and left out of your overall level rather than scored at all. If pronunciation is the thing you most need judged, a human rater is the only honest answer.
Structure
Control of grammatical structures and sentence patterns, from the basic ones you need constantly to the complex ones that carry reasoning: conditionals, past narration in sequence, passives, reported speech.
From a transcript: Judged well from text. The report quotes the actual clause rather than describing it, and separates an error that changes your meaning from one that merely marks you as a non-native speaker. Only the first holds a score down.
Vocabulary
Range and accuracy of vocabulary on operational topics, and the ability to paraphrase when the exact word does not arrive.
From a transcript: Judged well from text. A successful paraphrase is credited, not penalised: describing the cloth cone that shows the wind direction when windsock will not come is a skill, and it is scored as one. A missing word only counts against you when it stopped the message.
Fluency
Tempo, flow, connectors and discourse markers, and whether hesitation gets in the way of communicating.
From a transcript: Partly judged from text. Answer length against the question, subordination, restarts and abandoned clauses all survive into a transcript; the pauses do not. So this criterion is judged mainly on whether you developed an idea or stopped at one clause. Length is not the measure: three developed clauses score higher than ninety seconds of the same idea repeated.
Comprehension
Understanding what was asked, across routine and non-routine content, including when a question deliberately complicates things.
From a transcript: This is the criterion a transcript evidences best. An answer that addressed a different question than the one asked is direct evidence, and so is a correct answer to a deliberately awkward follow up.
Interactions
Managing the exchange: responding without delay, checking and confirming when unsure, and repairing a misunderstanding.
From a transcript: Judged well from text, and the strongest single signal is what you did when something was unclear. Asking the examiner to repeat or rephrase is good interaction behaviour and is scored as such. Guessing silently and answering the wrong question is the weak move.
How a level is produced
Four steps, and only one of them involves a language model at all.
- Your speech becomes text inside your own browser. The speech recognition your browser already provides does the work on your device. No audio is sent anywhere.
- At the end of the session, the text is sent once. The transcript goes with a record of how it was captured (how many times the microphone was opened, which words the recogniser missed on more than one separate answer, how much of what it heard you rewrote by hand before submitting) and about a dozen numbers per answer describing how it sounded (how long you spoke for, where your pauses fell, how fast the syllables came, how much background noise your microphone picked up). Those numbers are computed on your device from sound that is then thrown away.
- The model returns six scores and its evidence. Every score is required to come back with quotations of your own words and one concrete action, and an action line that is missing or too short to act on has the whole report rejected. It does not return your level, and it is not asked to.
- We compute the level ourselves. It is the lowest of the criteria the session could assess. Never an average, never rounded up.
The lowest, rather than the average, is the rule a real rating uses, and it is the honest one: if your structure falls apart under pressure, the fact that your vocabulary is excellent does not help the controller understand you. One weak criterion holds the whole result down.
The words could assess carry more weight than they look like they do. When the model says it had too little evidence to judge a criterion, that criterion is left out of the minimum altogether rather than counted at a placeholder level. We measured what happens without that rule: a candidate who typed every answer came back at overall level 3, and the two criteria dragging them down were pronunciation and fluency, which are exactly the two a typed session cannot evidence at all. Your report tells you how many of the six your level actually rests on, so you can weigh it.
Two more things our own code does rather than trusting the model with. It recomputes that minimum from the six scores instead of accepting an overall from the model. And if the summary or any of the six action lines claims you would pass or fail, that you are certified, that anything is guaranteed, or that this is an official rating, the whole report is thrown away rather than edited, and you are told the grading failed.
What we do when we cannot hear you properly
Before the interview starts, on the same screen as the Start button and above it, you read one fixed sentence out loud, and finishing the reading is what opens the interview. It is the same sentence for everybody, every time, and there is no reason not to print it here:
Golf Alpha Bravo Charlie Delta, good morning. We are passing flight level 190, descending to flight level 100 on heading 270. QNH 1013, and we have information Charlie. Request radar vectors for the approach on 121.5, and we would like to hold if the weather does not improve.
It is not scored as English and no part of your level comes from it. Because we already know what that sentence says, every word that comes back wrong is the machine being wrong. That is the whole point: it is the only measurement in the product taken against text we already know, and it is scored from the raw recogniser output rather than from anything you can edit.
Here is why it exists. We took one level 5 speaker's answers, put them through a real speech recogniser at a range of first-language accents, and graded the results with no change to the marking, the rubric or the model.
- About one word in ten transcribed wrongly, and that level 5 speaker was rated 4.
- About four words in ten wrong, and the same speaker was rated 3.
- Errors of two whole levels went from none at all on written text to 22.9 percent of runs once a real recogniser stood between the speaker and the grader.
The grader was not being unreasonable. It read a transcript in which one word in ten was wrong and correctly concluded that this candidate's grammar and vocabulary were not consistently controlled. It was being lied to about what the candidate said, and no amount of rewriting the marking instructions fixes a lie in the input. Note also who gets hurt: the damage falls hardest on the strongest speakers, because a candidate at level 3 cannot fall much further, while a candidate at 5 or 6 can lose two levels to a microphone.
So the rule is negative, and it is structural rather than a matter of judgement. If the recogniser reproduces less than 75 percent of that sentence, you are given no level at all. Not a lower one, not a hedged one, none. The session is handed back rather than charged, and there is deliberately no code path that could turn a poor recording into a poor result instead. Rating you on words you did not say would be marking you down for your microphone rather than for your English.
Between 75 and 90 percent you are rated and the report says plainly that some of your words reached the examiner wrongly. Above 90 percent you are rated and the report says how much of the sentence came through. Either way you are told the number rather than left to wonder. A reading below the floor stops there, before the session is counted, so it costs you nothing and you can pick up the headset and read it again.
Two situations where no reading happens, and we would rather name them than let you assume the floor always applied. If you are typing your answers there is nothing to calibrate, because no recogniser stands between what you wrote and what is rated, so you are offered the typed path instead of the sentence. And if you never attempt the reading at all and then speak anyway, we have no measurement of your recogniser, so we cannot apply the floor to you. The report then says outright that nothing measured how well your speech was being transcribed, rather than letting you read silence as reassurance.
What you cannot do is fail the reading and then leave it behind. A reading that came back below the floor is kept with the session, so if you start the interview anyway it still ends without a level and still returns the session to your account. Getting past the check is not the same as passing it, and only one of the two gets you a rating.
Now the honest part about that measurement, because publishing a finding without its limits is not publishing it. It was made on synthesised speech at a range of first-language accents, not on recordings of real pilots. What we rely on is the shape of the finding: that transcription damage costs whole levels, and that it costs the strongest speakers the most. The exact rates would move with real speakers. There is a second gap worth stating: the error rates above are for spontaneous speech, while the sentence you read is scripted and printed in front of you, which is an easier job for a recogniser. That makes this threshold err towards rating a session rather than refusing one.
What is not measured, and why
Pronunciation. A transcript cannot carry your accent, your stress, your rhythm or your intonation, and nothing in this product performs phoneme-level scoring of any kind today.
What a transcript can carry is the consequences: words that came through garbled, numbers that arrived wrong, places where you corrected yourself, points where the examiner had to rephrase, and words the recogniser missed on more than one separate answer. Those are signals in one direction only. Difficulty can support a lower level. The absence of difficulty supports nothing at all, because a recogniser coping is not a listener understanding you.
The practical consequence is the thing to take away, and it is easy to miss if we do not say it outright. On our own test transcripts pronunciation was marked as insufficient evidence on all 65 gradings, and was therefore excluded from every overall level in the sweep. Expect the same in a real session: your level will almost always rest on the other five criteria, your report will say so, and it will tell you what would actually produce evidence about your pronunciation. Leaving it out costs you nothing, because a criterion that is not counted cannot drag the minimum down, which is exactly why there is never a reason for the examiner to guess at it instead.
Fluency sits in between. Whether you developed an idea or stopped at one clause is visible in the text, and the pauses while you did it are measured on your device rather than read from the words. Comprehension is the criterion a transcript evidences best, because an answer that addressed a different question than the one asked is direct evidence.
Where your audio goes: nowhere
Your speech is turned into text by your own browser, on your own device, and the audio never leaves it. We never receive it, never store it, and never send it anywhere. There is no upload to switch off, because there is no upload. No recording of you exists, and we could not produce one if we were asked to.
Only text leaves your browser. At the end of the session it goes once to Anthropic's Claude API, which writes the report. Your name and your email address are not sent with it. Nothing goes to any airline, recruiter, training organisation or examiner, ever. Section 2.14 of the privacy policy is the detail: what is stored, who can read it, and the 12 month retention.
One browser has no built-in speech recognition: Firefox, on desktop and on Android. There, you type. A typed answer is read against the same criteria, with one honest exception that works in your favour: it is not evidence of pronunciation or fluency at all, so both are flagged and left out of your level rather than guessed at from your writing. If you would rather speak, your own device will dictate into the box for free, with the Windows key and H on Windows, two taps of Fn on a Mac, or the microphone key on a phone keyboard. Every other browser in common use, Chrome, Edge, Samsung Internet, and Safari on both Mac and iPhone, transcribes locally.
You can correct the text of an answer before you submit it. Two things you should know about doing so, because we would rather tell you than have you find out from the report. How much you rewrote is recorded and reaches the grader, as a measurement of the machine and you disagreeing rather than as a mark against you. And a heavily corrected answer is no longer evidence of how you sound, so correcting your way through a bad recording does not rescue the pronunciation reading, it removes it.
What we measured about our own marking, and published
We wrote 13 transcripts, each one built to a level we chose in advance, and graded every one of them five times on the same model that grades your session. That is 65 gradings of transcripts whose intended level we already knew.
- The intended level was matched exactly in 52 of 65 runs, 80.0 percent.
- It was one level out in 13 of 65 runs, 20.0 percent.
- It was never two or more levels out: 0 of 65 runs, 0.0 percent.
- Running the same transcript five times produced overall levels no more than one level apart, on 13 of 13 transcripts.
- Every single run quoted the candidate's own words as its evidence, on 13 of 13 transcripts.
Read those for what they are. They are measurements of our grader on our own test fixtures, not a claim about real candidates and not a claim about how a licensed rater would score anybody. They also describe one job only: the marking of a transcript that is already correct. What happens to those figures once a real microphone is in the loop is the section above, and it is the entire reason the calibration floor exists.
The session structure
Every session is the same shape, and the shape is enforced by our server rather than by the page you are looking at. There is no free chat and no way to keep talking past the structure.
- Part 1: interview. Questions about you, your background and your normal operating environment. Answer in full sentences and give reasons, not one word answers. About 5 minutes and 5 exchanges.
- Part 2: picture description. You will be shown one scene. Describe what you can see, explain why it matters, and say what you think happens next. The examiner will then ask you about it. About 5 minutes and 4 exchanges.
- Part 3: discussion. An open discussion about a non-routine situation. Expect follow up questions that push you to explain, compare and speculate. About 7 minutes and 6 exchanges.
That is 17 minutes of speaking time and 15 exchanges in total. Each part has its own clock. When the clock runs out the interview moves on, and when the last part ends it goes straight to the report.
The questions themselves are not written by a model. Every one of them, including the follow ups the examiner asks about your own answers, is written in advance by us and recorded in one consistent voice. What happens during the interview is that the examiner works out which subject you have just raised and picks the follow up that fits it, or, in the picture task, notices which part of the scene you have not described yet and asks about that.
The honest limit of that: the examiner can tell you it noticed the subject you raised, but it cannot repeat your own words back to you the way a human examiner would. What you get in exchange is an examiner that sounds the same from the first question to the last, works at the same speed on every connection, and cannot go quiet because your browser has no speech voice installed.
The limits, and why they exist
You get 25 mock interviews with your purchase, for the lifetime of the account. A session is counted the moment part 1 begins, whether or not you finish it, and the page tells you that before you start. Abandoning one still uses it.
The reason is cost, and it is worth being plain about it rather than calling it fair use. Every session spends real money on the language model that writes your report. We hold that under EUR 0.25 a session by capping the exchanges, capping the time, and stopping the session if it somehow costs more than expected. A one-time purchase with unlimited sessions would either be priced far higher or quietly degraded later, and we would rather tell you the real number now.
The 15 exchange cap is also a quality decision, not only a cost one. A longer interview does not produce a better report: the evidence for all six criteria is usually in place well before the clock runs out, and what a longer session mostly adds is repetition.
When you get no level
There are three of these, and they are not the same news.
The recording could not be rated. Your reading of the fixed sentence came back below 75 percent, so nothing you said afterwards can be judged fairly. The report says so, no scores are produced, and the session goes back onto your account automatically. You do not have to ask, and there is nothing to pay. A wired headset in a quiet room fixes this almost every time.
The session produced too little language to judge. If every criterion came back flagged for want of evidence, there is no minimum left to take, so the report shows no overall level and says what would produce one: answer more of the questions, and at greater length. That is the one case where the answer really is to speak more.
The grading step failed. Occasionally the report cannot be produced at all. When that happens the report says so and shows no scores rather than inventing them. It will never show you a low result that is really a technical failure. That case does use the session, so write to to2000bv@gmail.com and we will put it back.
The standing disclaimer
This is an AI-generated practice estimate, not an ICAO Language Proficiency assessment. It cannot award, predict, or guarantee any ICAO level or any test result. Only a licensed rater at an approved test centre can assess your ICAO level.