Active contractor · Mercor Generalist Expert · Les Artistes Artifacts
AI Lab Case · AI Writing Evaluation & Bilingual Editorial QA

A polished paragraph isn't a true one.

I evaluate AI-generated writing for what a fluency check misses: whether the instruction was actually followed, whether the claim is earned, and whether the sentence would survive a second read. Rated, sourced to evidence and written in the register a human grader can audit.

Why this judgment is real, not claimed

Digital Brand Strategy, JancisRobinson.comPublished, edited writing on a masthead with a fact-checking bar — not a personal blog.
Bilingual native ES · professional ENRioplatense Spanish as a first language, English as the working language of 15+ years in export and agency work.
Master of Wine candidate, Stage 2 · IWC judgingBlind, criteria-based scoring discipline — the backbone of a repeatable writing rubric.
Founder, Grand Crew StudioClient copy, proposals and brand voice systems written and edited end to end, in both languages.
Target rolesAI Response Evaluator · Bilingual AI Evaluator · Editorial QA Specialist
Core valueRubric-based writing evaluation with reproducible severity
Task typesRating, A/B comparison, error classification, ES-EN localization QA
Review modeVerdict → evidence → consequence, third person, 1–3 sentences
Case objective

Grade AI-generated writing beyond surface-level fluency.

For AI labs, this profile sits where editorial judgment, bilingual QA and instruction-following discipline overlap. The value isn't knowing what reads well. It's knowing whether a response actually did what was asked, whether its claims are backed, and whether that can be scored the same way twice.

1

Instruction following first

Identify exactly what was asked — rate, choose, correct or rewrite — before judging how it reads. Most severity errors start by skipping this step.

2

Evidence before opinion

Every judgment ties to a quoted phrase or a specific omission — never a general impression like "this feels off."

3

Calibrate severity

Minor, moderate, major or critical — and don't let a well-written sentence buy down a real instruction-following failure.

4

Write the verdict formula

Verdict, then evidence, then consequence — third person, 1–3 complete sentences, never "I think" or "this sounds AI-generated."

Evaluation framework

Eight lenses I use to score AI-generated writing.

Repeatable enough to guide annotation, flexible enough for rating tasks, A/B comparisons, error classification and Spanish-English localization QA.

01

Instruction following

Did it answer exactly what was asked — format, scale and constraints included?

02

Accuracy & evidence

Is every claim supported within the response, or does it assert something as certain with no backing?

03

Completeness

Does it cover everything the prompt required, or quietly drop a comparison, constraint or section?

04

Slop & buzzword detection

Business metaphors, filler adjectives and paragraphs that don't survive a second read — separate from legitimate direct instructions.

05

Register & localization

Does the Spanish sound natural, or grammatically correct but visibly translated?

06

Helpfulness

Could the user act on this without redoing the work themselves?

07

Severity calibration

Minor style note or major task failure — and not letting polish disguise the difference.

08

Rationale reproducibility

Could another grader reach the same score from my written rationale alone?

Sample AI output review · unedited

How I'd score a fluent but unsubstantiated AI-generated sentence.

AI output Market summary line

“Exports remain the anchor as digital, travel and payments add new momentum.”

  • Grammatically clean, confident register, reads like a market brief.
  • "Anchor" and "momentum" are consulting metaphors, not measurements.
  • No number, date or mechanism behind either claim.
  • Would pass a fluency check and fail an evidence check.
Rating 2 / 5

Verdict formula

The response is fluent but commercially vague, which makes it unusable as a market signal. "Anchor" and "momentum" describe direction without a number, date or mechanism behind either claim, so a reader cannot act on it or verify it. To improve it, name the metric behind "momentum" and quantify how exports compare to the other three categories.

  • Rule 1 check: no hard data present — the vagueness is real, not a false positive.
  • Rule 2 check: this isn't a direct instruction, so the buzzword flag stands.
  • Fix: replace "anchor" → "largest export category"; "momentum" → the actual growth rate.
Line-level example

Before and after: from confident-sounding filler to a verifiable claim.

Generated · vague / empty

“Automation increases efficiency.”

Active voice, short sentence, correct grammar — and still says nothing a reader could act on or check.

Grading note: active voice is not immunity. No number, date or mechanism → slop, whatever the tone.

Fixed · specific / verifiable
What the fix adds

“Automating intake cut onboarding time from 12 days to 4.”

Subject, action, number and mechanism — a claim a reader could verify against a source.

Grading note: same length, same register — the difference is entirely evidentiary, not stylistic.

Role requirements
Onboarding status · active, all gates cleared

What Mercor required for this role, and where it stands.

Generalist Expert on Les Artistes Artifacts, USD 50/h, at-will, since 25-Aug-2026 — every onboarding step below is complete and the contract is running.

Done

Contract & background check

Offer Letter, Worker Agreement and Invention Assignment signed. Background check (criminal record + identity, via Zincwork) completed and clear 27-Aug-2026.

contractcompliance clear
Done

Evaluation courses

Completed the project's evaluation training: deck outline, style matching, template creation, complete-the-deck — rubric format with slide reference, standard-linked rationale, third person, three sentences.

training complete
Active

Time tracking & task tooling

Insightful (Workpuls) running for mandatory time capture; Feather provisioned for task access. Separate, AI-free macOS user for the contracted work itself, per the Worker Agreement.

InsightfulFeather
Rule

No third-party AI on the work itself

The Worker Agreement bars inputting confidential project content into outside AI tools, or using AI output in Developed IP. This calibration practice runs on public, synthetic material only — never the actual project content.

confidentiality
Ready

Bilingual editorial discipline

Native Rioplatense Spanish, professional English, published editorial work at JancisRobinson.com — the same register-and-accuracy judgment the Spanish-English QA rubric asks for.

ES-EN QA
Tracked

Calibration accuracy

79% on a 29-case reference calibration round, with a named blind spot (buzzword and vague-wording misses) and a working correction rule for each.

self-audited

Status

Active contractor, milestone M2 in progress — the first-project objective for September. No access blockers on this contract; the September goal is sustained volume, not onboarding.

Evidence map

Why this profile is believable from a real career.

The case connects public-facing bilingual writing and editorial work with a private AI-evaluation layer. It doesn't pretend to be a pure linguist — it positions the rubric-disciplined, bilingual editorial judgment labs are short of.

Published, edited writing

JancisRobinson.com and Wine-Searcher work — writing that answers to an editor and a fact-checking bar, not self-published content.

editorial judgmentcontent clarityfact-sensitive editing

Bilingual business writing

15+ years writing and editing commercial content in both languages — export sales, agency client copy, brand voice systems.

Spanish-English translation evaluationregister matchingtone and register

AI evaluation layer

This case converts that experience into AI-lab vocabulary: rating, error classification, localization QA and rubric-based feedback.

AI output evaluationrubric-based feedbackmodel output review
Focus areas

The AI writing work I'm strongest at reviewing.

Mercor-facing keywords, copied exactly from the archetype bank — Bilingual Spanish-English Editorial QA archetype plus the transversal AI-evaluation core.

AI Response Evaluator Bilingual AI Evaluator Editorial QA Specialist AI Writing Evaluator Spanish AI Translation and Evaluation Expert LLM Response Quality Analyst AI output evaluation instruction following error classification hallucination detection rubric-based feedback localization QA bilingual review editorial judgment model output review response ranking quality scoring severity classification written rationale content quality analysis
Flor Gómez, writer and bilingual editorial reviewer
Portrait · 4:5
Work_AiLabClosing_retrato.jpg
The reviewer

The judgment behind the rating.

I'm Flor Gómez — bilingual writer, founder of Grand Crew Studio and Master of Wine Stage 2 candidate. I've written and edited for JancisRobinson.com, judged blind at Decanter and IWC, and I review AI writing the same way: instruction first, evidence before opinion, verdict with a number attached.

Flor Gómez · Barcelona · English & Spanish · flor-gomez.com

For Mercor and similar AI writing / editorial QA projects

You've seen the sample rating. If that's the standard you need, contact me.

AI writing evaluation, Spanish-English editorial QA, instruction-following review and rubric-based feedback. Tell me the project and I'll reply with fit, scope and rate within one business day.

Prefer to write? Use the contact form · Related case: Creative Director / Art Director — AI Deck Grading (Ethos) · Portfolio: flor-gomez.com

Start here

Contact me.

Goes straight to my inbox — no middlemen. Tell me what you need, paste a link if you have one, and I'll reply within one business day.

What's this about?

I reply within one business day.

Got it — thank you.

Your message is on its way to me. I'll reply within one business day. If it's time-sensitive, you can also book a 15-min call.