AI-Native Engineering Studio

We build production AI
that ships

Agentic pipelines, LLM platforms, and the evals behind them. Recent work includes writing-quality evals for OpenAI and an agent platform Signet Jewelers stood up in three months.

years shipping production software
10+
years shipping production software
deliverables scored for OpenAI evals
60+
deliverables scored for OpenAI evals
cut in Mayo Clinic validation time
94%
cut in Mayo Clinic validation time
AI-writing patterns in unslopify
27
AI-writing patterns in unslopify
OpenAIMayo ClinicSignet JewelersJP Morgan ChaseCACI / US ArmyRivianCollege BoardTicketmasterEtsyVeterans AffairsBayer

Our Work

Client platforms and our own products, with the results each one shipped.

Writing-Quality & Presentation Evals

OpenAI, via Mercor

Expert evaluation on OpenAI's Feather platform: evals that catch AI slop in model-generated business writing and slide decks. Every score lands on a calibrated 1 to 7 rubric with a written rationale, and the whole workflow runs as a pipeline with deterministic gates before anything ships.

EvalsLLM-as-judgeRubric designPythonHuman-in-the-loop

Key Results

  • 60+ deliverables scored across writing and presentation campaigns
  • Deterministic audit scripts gate every rationale before submission
  • Two fresh-context model judges re-read each draft cold
  • Reviewer sendbacks folded back into the grading rules

AEM Content Validator

Mayo Clinic

Production LangGraph pipeline with 5 parallel agents and a zero-cost triage router that settles easy cases without calling a model. Validates articles against 34 rules loaded from Neo4j at runtime with a JSON fallback, with retrieval on PGVector measured by RAGAS before rollout.

LangGraphNeo4jPGVectorRAGASFastAPINext.jsSSE

Key Results

  • 94% reduction in validation time (2 hours to under 10 minutes)
  • 5 parallel agents with zero-cost triage router
  • RAGAS-measured faithfulness and recall before rollout
  • Reviewer dashboard with SSE streaming and approval gates

AI Foundry Platform

Signet Jewelers

Enterprise AI platform for the world's largest diamond jewelry retailer, shipped in three months: a React agent studio on AWS CDK with a RAG hub internal teams point at their own documents. Bedrock and AgentCore streaming chat runs behind a custom Vercel AI SDK provider with SSE heartbeats.

AWS BedrockAgentCoreReactTypeScriptAWS CDKMCP

Key Results

  • Platform live in 3 months with self-serve team onboarding
  • MCP server, model catalog, and per-agent spend tracking
  • Rebuilt memory reconciliation to survive CloudFront timeouts
  • 3-agent copilot took PR review from days to hours

unslopify

V2 Software product, open source

Learn more

A linter for prose that names what a model got wrong: 27 pattern types across formula, substance, wording, and structure. The audit core is deterministic Python with no model calls, so it runs in CI, and a cross-document phrase bank catches a writer or an agent reusing its own stock phrasing.

PythonPydanticPyPICLIAgent skill

Key Results

  • 27 named patterns in four categories, each with severities
  • Deterministic core, zero model calls, CI-friendly exit codes
  • Cross-document phrase bank stops repeated stock phrasing
  • Open source under MIT on GitHub and PyPI

JobHunter Agent

V2 Software product

Visit

Full-stack autonomous job hunting platform with LangGraph agentic pipelines, browser automation via Patchright, resume parsing, and real-time SSE streaming. Automates job discovery, application submission, and follow-up across multiple job boards.

LangGraphFastAPINext.jsPatchrightPostgreSQLRedisRailway

Key Results

  • Autonomous multi-agent pipeline for end-to-end job applications
  • Real-time SSE streaming with human-in-the-loop approval gates
  • Stripe-integrated SaaS with tiered pricing
  • Browser automation with anti-detection for major job boards

ML Visualization Dashboards

CACI / US Army

Classified ML visualization dashboards and secure web applications for US Army intelligence analysts. Interactive data tools for mission-critical decision-making on DoD networks, delivered with an active Secret Clearance.

ReactPythonAWS GovCloudD3.jsPostgreSQL

Key Results

  • Classified ML dashboards for US Army intelligence analysts
  • Interactive visualization for mission-critical decisions
  • Hardened applications deployed on DoD networks
  • Full-stack delivery on AWS GovCloud infrastructure

Vehicle Parts Sourcing Platform

Rivian Automotive

Internal parts sourcing platform used by Rivian's supply chain team to manage component procurement for R1T and R1S vehicle production. Real-time supplier tracking, cost analysis, and procurement workflows through the production ramp.

ReactAWS AmplifyGraphQLDynamoDBTypeScript

Key Results

  • Managed component procurement for R1T and R1S production
  • Real-time supplier tracking and cost analysis dashboards
  • Procurement workflows supporting the production ramp
  • Built on AWS Amplify with GraphQL and DynamoDB

About the Founder

Fareez Ahmed is the CEO of V2 Software LLC, a senior AI engineer with over ten years across full-stack web applications, backend services, and production LLM systems.

Recent work: writing evals for OpenAI and the Mayo Clinic validator that cut review time by 94 percent. We also build and maintain unslopify, an open-source writing-quality linter.

MA in International Affairs and Economics from Columbia University. Active Secret Clearance. Based in San Francisco and Dallas, working remote. Speaks English, French, Arabic, and Bengali.

Client Experience

OpenAI (via Mercor)
Writing-Quality & Presentation Evals, Feather
Mayo Clinic
AEM Content Validator, LangGraph Pipeline
Signet Jewelers
AI Foundry Platform, Bedrock & AgentCore
JP Morgan Chase
React/Redux Microfrontends, AWS Serverless
CACI / US Army
ML Visualizations, Secure Applications
Rivian Automotive
Vehicle Parts Sourcing, AWS Amplify
College Board
Student Portal, AWS Serverless at Scale
Live Nation
Ticketmaster, Event Discovery at Scale

Tech Stack

The tools we ship with, grouped by layer.

AI / LLM

LangGraphLangChainLlamaIndexBedrock AgentCoreOpenAI APIAnthropic APIVertex AILiteLLMRAGASLangSmithCrewAIMCP Protocol

Backend

Node.jsNestJSExpress.jsPythonFastAPIFlaskRuby on RailsJava

Frontend

ReactNext.jsTypeScriptJavaScriptReduxTailwind CSSAngularD3.js

Databricks

JobsUnity CatalogDelta LakeAsset BundlesSQL WarehousesModel ServingAI GatewayDatabricks AppsLakebase PostgresMLflow 3

Cloud / Infra

AWS LambdaS3SQSStep FunctionsAWS CDKTerraformTerragruntKubernetesArgoCDIstioGCP Vertex AIAzure MLRailway

Databases

PostgreSQLDynamoDBMongoDBNeo4jPGVectorRedisPinecone

Testing & DevOps

JestPytestPlaywrightCypressStorybookGitHub ActionsGitLab CICircleCI

Let's Build

Need a production AI system, agentic pipeline, or full-stack platform? Let's talk about what you're building.

Location
San Francisco, CA / Dallas, TX (Remote)