A polished paragraph isn't a true one.
I evaluate AI-generated writing for what a fluency check misses: whether the instruction was actually followed, whether the claim is earned, and whether the sentence would survive a second read. Rated, sourced to evidence and written in the register a human grader can audit.
Grade AI-generated writing beyond surface-level fluency.
For AI labs, this profile sits where editorial judgment, bilingual QA and instruction-following discipline overlap. The value isn't knowing what reads well. It's knowing whether a response actually did what was asked, whether its claims are backed, and whether that can be scored the same way twice.
Instruction following first
Identify exactly what was asked — rate, choose, correct or rewrite — before judging how it reads. Most severity errors start by skipping this step.
Evidence before opinion
Every judgment ties to a quoted phrase or a specific omission — never a general impression like "this feels off."
Calibrate severity
Minor, moderate, major or critical — and don't let a well-written sentence buy down a real instruction-following failure.
Write the verdict formula
Verdict, then evidence, then consequence — third person, 1–3 complete sentences, never "I think" or "this sounds AI-generated."
Eight lenses I use to score AI-generated writing.
Repeatable enough to guide annotation, flexible enough for rating tasks, A/B comparisons, error classification and Spanish-English localization QA.
Instruction following
Did it answer exactly what was asked — format, scale and constraints included?
Accuracy & evidence
Is every claim supported within the response, or does it assert something as certain with no backing?
Completeness
Does it cover everything the prompt required, or quietly drop a comparison, constraint or section?
Slop & buzzword detection
Business metaphors, filler adjectives and paragraphs that don't survive a second read — separate from legitimate direct instructions.
Register & localization
Does the Spanish sound natural, or grammatically correct but visibly translated?
Helpfulness
Could the user act on this without redoing the work themselves?
Severity calibration
Minor style note or major task failure — and not letting polish disguise the difference.
Rationale reproducibility
Could another grader reach the same score from my written rationale alone?
How I'd score a fluent but unsubstantiated AI-generated sentence.
“Exports remain the anchor as digital, travel and payments add new momentum.”
- Grammatically clean, confident register, reads like a market brief.
- "Anchor" and "momentum" are consulting metaphors, not measurements.
- No number, date or mechanism behind either claim.
- Would pass a fluency check and fail an evidence check.
Verdict formula
The response is fluent but commercially vague, which makes it unusable as a market signal. "Anchor" and "momentum" describe direction without a number, date or mechanism behind either claim, so a reader cannot act on it or verify it. To improve it, name the metric behind "momentum" and quantify how exports compare to the other three categories.
- Rule 1 check: no hard data present — the vagueness is real, not a false positive.
- Rule 2 check: this isn't a direct instruction, so the buzzword flag stands.
- Fix: replace "anchor" → "largest export category"; "momentum" → the actual growth rate.
Before and after: from confident-sounding filler to a verifiable claim.
What Mercor required for this role, and where it stands.
Generalist Expert on Les Artistes Artifacts, USD 50/h, at-will, since 25-Aug-2026 — every onboarding step below is complete and the contract is running.
Contract & background check
Offer Letter, Worker Agreement and Invention Assignment signed. Background check (criminal record + identity, via Zincwork) completed and clear 27-Aug-2026.
Evaluation courses
Completed the project's evaluation training: deck outline, style matching, template creation, complete-the-deck — rubric format with slide reference, standard-linked rationale, third person, three sentences.
Time tracking & task tooling
Insightful (Workpuls) running for mandatory time capture; Feather provisioned for task access. Separate, AI-free macOS user for the contracted work itself, per the Worker Agreement.
No third-party AI on the work itself
The Worker Agreement bars inputting confidential project content into outside AI tools, or using AI output in Developed IP. This calibration practice runs on public, synthetic material only — never the actual project content.
Bilingual editorial discipline
Native Rioplatense Spanish, professional English, published editorial work at JancisRobinson.com — the same register-and-accuracy judgment the Spanish-English QA rubric asks for.
Calibration accuracy
79% on a 29-case reference calibration round, with a named blind spot (buzzword and vague-wording misses) and a working correction rule for each.
Status
Active contractor, milestone M2 in progress — the first-project objective for September. No access blockers on this contract; the September goal is sustained volume, not onboarding.
Why this profile is believable from a real career.
The case connects public-facing bilingual writing and editorial work with a private AI-evaluation layer. It doesn't pretend to be a pure linguist — it positions the rubric-disciplined, bilingual editorial judgment labs are short of.
Published, edited writing
JancisRobinson.com and Wine-Searcher work — writing that answers to an editor and a fact-checking bar, not self-published content.
Bilingual business writing
15+ years writing and editing commercial content in both languages — export sales, agency client copy, brand voice systems.
AI evaluation layer
This case converts that experience into AI-lab vocabulary: rating, error classification, localization QA and rubric-based feedback.
The AI writing work I'm strongest at reviewing.
Mercor-facing keywords, copied exactly from the archetype bank — Bilingual Spanish-English Editorial QA archetype plus the transversal AI-evaluation core.
Work_AiLabClosing_retrato.jpg
The judgment behind the rating.
I'm Flor Gómez — bilingual writer, founder of Grand Crew Studio and Master of Wine Stage 2 candidate. I've written and edited for JancisRobinson.com, judged blind at Decanter and IWC, and I review AI writing the same way: instruction first, evidence before opinion, verdict with a number attached.
Flor Gómez · Barcelona · English & Spanish · flor-gomez.com
You've seen the sample rating. If that's the standard you need, contact me.
AI writing evaluation, Spanish-English editorial QA, instruction-following review and rubric-based feedback. Tell me the project and I'll reply with fit, scope and rate within one business day.
Prefer to write? Use the contact form · Related case: Creative Director / Art Director — AI Deck Grading (Ethos) · Portfolio: flor-gomez.com
Contact me.
Goes straight to my inbox — no middlemen. Tell me what you need, paste a link if you have one, and I'll reply within one business day.
Got it — thank you.
Your message is on its way to me. I'll reply within one business day. If it's time-sensitive, you can also book a 15-min call.