Writing-Quality & Presentation Evals
OpenAI (via Mercor)
- Expert evaluator on OpenAI's Feather platform, building evals that detect AI slop in model-generated business writing and presentation decks.
- Scored 60+ deliverables on a calibrated 1 to 7 rubric covering writing quality, organization and storytelling, comprehensiveness, and groundedness, each with a written rationale that held up under reviewer audit.
- Rebuilt the evaluation workflow as a pipeline: deterministic audit scripts, ranking-spread and phrase-uniqueness checks, and two fresh-context model judges before every submission.