We build production AI
that ships
Agentic pipelines, LLM platforms, and the evals behind them. Recent work includes writing-quality evals for OpenAI and an agent platform Signet Jewelers stood up in three months.
- years shipping production software
- 10+years shipping production software
- deliverables scored for OpenAI evals
- 60+deliverables scored for OpenAI evals
- cut in Mayo Clinic validation time
- 94%cut in Mayo Clinic validation time
- AI-writing patterns in unslopify
- 27AI-writing patterns in unslopify
Our Work
Client platforms and our own products, with the results each one shipped.
Writing-Quality & Presentation Evals
OpenAI, via Mercor
Expert evaluation on OpenAI's Feather platform: evals that catch AI slop in model-generated business writing and slide decks. Every score lands on a calibrated 1 to 7 rubric with a written rationale, and the whole workflow runs as a pipeline with deterministic gates before anything ships.
Key Results
- 60+ deliverables scored across writing and presentation campaigns
- Deterministic audit scripts gate every rationale before submission
- Two fresh-context model judges re-read each draft cold
- Reviewer sendbacks folded back into the grading rules
AEM Content Validator
Mayo Clinic
Production LangGraph pipeline with 5 parallel agents and a zero-cost triage router that settles easy cases without calling a model. Validates articles against 34 rules loaded from Neo4j at runtime with a JSON fallback, with retrieval on PGVector measured by RAGAS before rollout.
Key Results
- 94% reduction in validation time (2 hours to under 10 minutes)
- 5 parallel agents with zero-cost triage router
- RAGAS-measured faithfulness and recall before rollout
- Reviewer dashboard with SSE streaming and approval gates
AI Foundry Platform
Signet Jewelers
Enterprise AI platform for the world's largest diamond jewelry retailer, shipped in three months: a React agent studio on AWS CDK with a RAG hub internal teams point at their own documents. Bedrock and AgentCore streaming chat runs behind a custom Vercel AI SDK provider with SSE heartbeats.
Key Results
- Platform live in 3 months with self-serve team onboarding
- MCP server, model catalog, and per-agent spend tracking
- Rebuilt memory reconciliation to survive CloudFront timeouts
- 3-agent copilot took PR review from days to hours
unslopify
V2 Software product, open source
A linter for prose that names what a model got wrong: 27 pattern types across formula, substance, wording, and structure. The audit core is deterministic Python with no model calls, so it runs in CI, and a cross-document phrase bank catches a writer or an agent reusing its own stock phrasing.
Key Results
- 27 named patterns in four categories, each with severities
- Deterministic core, zero model calls, CI-friendly exit codes
- Cross-document phrase bank stops repeated stock phrasing
- Open source under MIT on GitHub and PyPI
JobHunter Agent
V2 Software product
Full-stack autonomous job hunting platform with LangGraph agentic pipelines, browser automation via Patchright, resume parsing, and real-time SSE streaming. Automates job discovery, application submission, and follow-up across multiple job boards.
Key Results
- Autonomous multi-agent pipeline for end-to-end job applications
- Real-time SSE streaming with human-in-the-loop approval gates
- Stripe-integrated SaaS with tiered pricing
- Browser automation with anti-detection for major job boards
ML Visualization Dashboards
CACI / US Army
Classified ML visualization dashboards and secure web applications for US Army intelligence analysts. Interactive data tools for mission-critical decision-making on DoD networks, delivered with an active Secret Clearance.
Key Results
- Classified ML dashboards for US Army intelligence analysts
- Interactive visualization for mission-critical decisions
- Hardened applications deployed on DoD networks
- Full-stack delivery on AWS GovCloud infrastructure
Vehicle Parts Sourcing Platform
Rivian Automotive
Internal parts sourcing platform used by Rivian's supply chain team to manage component procurement for R1T and R1S vehicle production. Real-time supplier tracking, cost analysis, and procurement workflows through the production ramp.
Key Results
- Managed component procurement for R1T and R1S production
- Real-time supplier tracking and cost analysis dashboards
- Procurement workflows supporting the production ramp
- Built on AWS Amplify with GraphQL and DynamoDB
About the Founder
Fareez Ahmed is the CEO of V2 Software LLC, a senior AI engineer with over ten years across full-stack web applications, backend services, and production LLM systems.
Recent work: writing evals for OpenAI and the Mayo Clinic validator that cut review time by 94 percent. We also build and maintain unslopify, an open-source writing-quality linter.
MA in International Affairs and Economics from Columbia University. Active Secret Clearance. Based in San Francisco and Dallas, working remote. Speaks English, French, Arabic, and Bengali.
Client Experience
Tech Stack
The tools we ship with, grouped by layer.
AI / LLM
Backend
Frontend
Databricks
Cloud / Infra
Databases
Testing & DevOps
Let's Build
Need a production AI system, agentic pipeline, or full-stack platform? Let's talk about what you're building.