Embedded production engineering

Production engineering for agentic AI systems.

We help AI product and startup engineering teams diagnose and fix production agent failures across evals, memory, tool use, state, observability, and human control.

Illustration of a request moving through connected panels to an approval checkpoint
  • AI product + startup teamsAlready building or shipping agentic systems
  • One failureReproduced and isolated before the fix
  • Tests + evalsRegression protection stays with your team
  • Reproduce the real failure
  • Fix the system, not the symptom
  • Leave tests, evals, and instrumentation behind
What we solve

When the happy path stops being enough.

  • Failures that will not reproduce

    Turn an intermittent production trace into a controlled case the team can rerun and inspect.

  • Eval gaps and regressions

    Build the benchmark that exposes the failure, then keep it as a release gate after the fix.

  • Memory and state failures

    Trace stale context, broken state transitions, and long-running sessions that drift from intent.

  • Unreliable tool execution

    Debug tool selection, malformed arguments, retries, side effects, and partial execution across real integrations.

  • Missing observability

    Instrument decisions, tool calls, state changes, and outcomes so the next failure has evidence attached.

  • Human control in production

    Design approval, escalation, and takeover paths for workflows where autonomous action is not enough.

How we work

Failure → reproduction → eval → fix → verification → regression protection.

We embed with the engineering team long enough to reproduce the failure and strengthen verification around it—not just patch the visible symptom.

  1. Failure

    Start with one concrete production behavior that should not have happened.

  2. Reproduction

    Reduce logs, traces, state, and inputs to a case we can run deliberately.

  3. Benchmark / eval

    Turn the case into an executable definition of correct behavior.

  4. Fix

    Repair the model, prompt, tool, state, or control boundary causing the failure.

  5. Verification

    Run the fix against the failing case and the surrounding behavior it could affect.

  6. Regression protection

    Leave the team with tests, evals, and instrumentation that catch recurrence.

Technical proof

Projects, experiments, and open-source work.

GraphKeeper

Open-source project

Grounded, auditable, Git-backed memory for coding agents.

  • Stores durable claims beside the code with links to the run and exact evidence that produced them.
  • Validates schema, provenance, append-only history, and superseding corrections through a CLI and Git hooks.
  • TypeScript
  • Git
  • Validation
  • Agent memory
Explore GraphKeeper

Agent Replay

Open-source experiment

A flight recorder and time-travel debugger for AI agents.

  • Records agent steps so a failure can be inspected and replayed instead of inferred from the final output.
  • Frames debugging around forking the failed run, proving the fix, and protecting the behavior from regression.
  • Python
  • Tracing
  • Replay
  • Agent debugging
View the repository

Lead qualification agent

Technical case study

A historical agent project built around inspectable state transitions and human takeover.

  • Webhook-driven finite state machine with one inference per turn, making every transition inspectable.
  • Business context stays separate from the engine, and an owner reply makes the agent stand down.
  • FastAPI
  • PostgreSQL
  • OpenRouter
  • Railway
Read the case study

Prospector

Open-source project

A CLI pipeline that turns a raw company list into sourced drafts for human approval.

  • Keeps confidence scoring, signal classification, and validation deterministic and tested.
  • Rejects unsourced names and unsupported claims, with a human approval gate before every send.
  • Python
  • CLI
  • Validation
  • Human approval
View the repository
Who you work with

Anas Butt

Anas Butt, founder of Omniveer

Karachi, Pakistan — working with engineering teams internationally

I work spec-first: reproduce the failure, write down the constraints and expected behavior, then fix against an executable benchmark. The goal is a system your team can debug and control after the engagement ends.

Bring us one agent failure

Start with the behavior that broke.

Send what the agent was expected to do, what actually happened, and any logs, traces, or reproduction details you already have.

Email the failure directly anas@omniveer.com
Or walk through it together Book a 20-minute technical intro