There is no score, and that is the central decision in the product. A number would have to come from somewhere, and there is no external scale for how good an answer about V1 was, so any figure invented here would be a figure only this site believed, varying between runs, presented to a candidate as a measurement.
What replaces it is coverage against a written model answer. Every question in both banks carries its key points, written by a person before anybody was ever asked it. The report says which of those points your answer reached and which it did not, quotes you, and names anything you said that was actually wrong. That is a checklist rather than a verdict, so two runs agree far more often than a graded one would, and where they disagree you can read the key point and decide for yourself.
The distinction the report works hardest to keep is between a missed point and a correction, because a pilot needs to know which one they did. A missed point is something you did not say, and it goes in the coverage. A correction is something you said that is wrong, and it is stated in one sentence with what is actually the case. An empty corrections list is normal and common, and the grader is told not to manufacture one to look thorough.
The evidence is a verbatim quotation from you, one sentence or two, copied from the transcript. Never the examiner, never a paraphrase presented as a quotation, and never a quotation assembled out of separated fragments. Where an answer contains no sentence worth quoting, it quotes the most representative one anyway, because reading your own thin sentence back is itself the feedback.
Each question ends on one action line, addressed to you, and it has to be actionable in preparation rather than in an aircraft. Revise the relationship between load factor and stall speed until you can state it without thinking is an action. Be more confident is not. Know your systems is not.
All of that is instruction to a model, and an instruction is not a guarantee. So the guarantee is code. Every report passes a normaliser before anyone can see it, and it refuses rather than repairs, because a repair would leave the surrounding sentence saying something the model meant and this product did not.
Start with the part that is structural rather than watchful: there is no field for a score anywhere in the shape the model has to answer in. It carries coverage, corrections, a quotation and an action, and nothing else, so a mark is not something the report could render even if the model wanted one. What is left is prose, and the prose is refused outright if it says you will pass or fail, that anything is guaranteed or certified, that your score, mark or grade is anything, that you passed or failed this interview, or that you are ready for the real one. That second half is a list of phrasings and not a reading of meaning, which is the honest limit of it. It drops feedback on a question that was never asked, which is what a model does when it invents one. It ignores a key point index that does not exist, which would otherwise render as a blank bullet. And it recomputes the missed list from the covered list on the server rather than reading it back from the model, so the two cannot contradict each other and the count underneath is arithmetic this site did.
A skipped question gets no feedback at all. Not a gentle note, not a placeholder: nothing. A skip is a smaller sample, never a wrong answer, and inventing feedback for silence would be the worst thing the report could do.