Skip to content
Back to homeTrust & grading

Last updated · July 2026

How we grade

Vantage grades Writing and Speaking with an AI examiner, not a human panel. We think you should be able to see exactly how that works before you trust the score. This page explains the pipeline in plain language, shows a worked example, and is upfront about where the AI can be wrong.

1. What happens when you submit

When you submit a Writing response or a Speaking recording, Vantage sends it to Google’s Gemini model together with a system prompt that instructs it to act as “an experienced CEFR examiner” and score the response using the same four criteria a human examiner would use, on the standard CEFR 1–6 scale (1 = A1 … 6 = C2). The model returns a structured result — a score for each criterion, an overall CEFR level, short examiner-style feedback, and (for Writing) inline annotations that quote the exact phrase they refer to. That result is what you see on your Results page.

There is no separate “grading algorithm” behind the scenes beyond this — the AI model is the grader. We don’t blend in a rules-based score or adjust the model’s numbers before showing them to you.

2. The four criteria

Writing is scored against the same four criteria examiners use for CEFR-aligned writing assessment:

  • TR

    Task Response — how fully and relevantly your response addresses the prompt.

  • CC

    Coherence & Cohesion — paragraphing, logical progression, and how well ideas are linked together.

  • LR

    Lexical Resource — the range, precision, and appropriateness of the vocabulary you use.

  • GR

    Grammatical Range & Accuracy — the variety of grammatical structures you use and how accurate they are.

Speaking uses an equivalent four-criterion set, swapping paragraphing for delivery: Fluency (FL), Pronunciation (PR), Lexical Resource (LR), and Grammatical Range & Accuracy (GR). The AI examiner listens to your recording directly — audio in, transcript and scores out — there’s no separate speech-to-text step you can’t see.

3. A worked example

Below is an illustrative example we wrote for this page — not a real candidate’s submission — to show how a sentence maps to a specific score.

Illustrative example · not a real submission

“Although many people think social media is only harmful, I believe it depends on how it is used, and there is many benefit for students who use it to study together.”

CC — “Although many people think… I believe it depends” is a cohesive concession-then-position move — the kind of clause-linking control the CEFR B2 Coherence descriptor (“can use a limited number of cohesive devices to link utterances into clear, coherent discourse”) is looking for.

GR — “there is many benefit” has a subject-verb agreement error and a missing plural, which is exactly the kind of slip that keeps a response from clearing the CEFR B2 Grammatical Accuracy bar (“good grammatical control; occasional slips”) — GR is marked down here even though TR and CC are strong.

4. Human review

AI feedback becomes visible on your Results page as soon as grading finishes — you don’t wait for a human to approve it first. Separately, Vantage staff review completed AI feedback through an internal queue and can approve, edit, or reject entries as a quality-control check on the model. That review happens independently of — not before — what you see, so treat the score you receive as the AI examiner’s live output, not a human-verified final mark.

5. Be honest with yourself about the score

Vantage’s AI-estimated CEFR level is a strong practice signal, not an official result. Your real Multilevel exam score may differ by roughly half a band (±1 sublevel) from what the AI reports here. Use it to track trends over time and target your weakest skill — not as a guarantee of your exam outcome.

We don’t publish an accuracy percentage or a “matches official results X% of the time” claim, because we don’t have that data yet — and we’d rather tell you that plainly than make up a number.

6. Questions

If a score looks wrong or a comment doesn’t make sense, Message us on Telegram with the test ID from your Results page — we read every report.

See it on your own writing — start a free mock.