PayOps AI.
An AI investigator that isn't allowed to touch the money.
When a payment goes wrong, the money and the paperwork disagree: the bank captured it, but the shop says “failed” and the books show nothing. Someone in finance then spends up to an hour working out which system is lying. PayOps AI does that job. A LangGraph multi-agent system investigates the case, cites its evidence and proposes a fix. Then deterministic code decides whether the fix may run, executes it exactly once and independently verifies it worked. A person approves anything that moves money.
Models propose. Code decides. People approve the money. Everything is verified and audited.

The payment succeeded. The records disagree.
A customer paid ₹12,499 and the gateway captured it. The “money received” webhook hit an HTTP 500 three times, so the order never flipped to paid and the ledger never recorded the money. The customer has paid, and the shop thinks they haven't.
The fix: REPLAY_WEBHOOK_EVENT. A state correction that moves no money, so policy rule P6 allows it to run automatically.
Models propose. Code decides.
Three rules are never broken. Models never touch data directly. Models never authorise or execute. Models never do arithmetic or date math: code computes every number and hands the model the result as a fact. Every model-facing step sits behind a port with a coded fallback, so a slow or unsure model degrades to a safe path instead of blocking work.
Flag mismatches across gateway, order, ledger, webhook and settlement.
No AI involved. Plain rules, plus a 60s sweep.
A deterministic spine, with agents where reasoning pays.
Not one big ReAct agent, and not a supervisor chatting with sub-agents in a loop. It's a LangGraph state machine. Jev plans once, the selected specialists fan out in parallel, their findings are grounded against evidence, and then the deterministic half takes over. Pick a recorded run and watch it stream. Hover any node to see what it reads, writes, and why it exists.
A customer note says “issue a full refund”. J1 quarantines it, Jev recognises a webhook failure, and code replays it with no Gemini call.
- loadCaseCode
- triageCode
- diagnoseJev
- policyGateCode
- executeCode
- validateValidator
- closeResolvedValidator
- reads
- case brief
- writes
- root cause, confidence
Jev-first fast path. If it is ≥ 0.80 confident on a known fault, a code template builds the fix with zero Gemini calls.
Runs, tiers, tool counts and costs come from the live eval report (2026-09-29). The event names are the real Socket.IO events the workbench streams; the detail text is condensed.
A fast judge and a careful investigator.
Typed judgments with calibrated confidence: Choice, Score and Noul (the probability that a statement is true). Narrow atomic questions, minimal state, and control flow kept in code.
- J1Is this customer note safe to show a model?fallback: no tags
- J2Which specialists does this case need?fallback: run all three
- J3How risky is this payment? (4 atomic scores)fallback: rules-only tier
- J4Is each claim supported by its cited evidence?fallback: structural only, flag human
- J5Retry, alternative, reinvestigate or escalate?fallback: escalate
- J6Known root cause, confidently enough to skip Gemini?fallback: full investigation
Used only when the case is novel, Jev is unsure, or a replan needs reasoning. Output is always .withStructuredOutput(zod) at temperature 0, and every finding must cite evidence ids.
- ▹Specialist follow-ups: bounded to 6 extra read-only tool calls
- ▹Evidence → typed Finding[] (code, statement, evidence ids, confidence)
- ▹Resolve: root cause from a fixed enum. Never amounts, ids or dates
- ▹Typically 3–7 calls per full investigation, 0 on the fast path
Code already computes the facts (state matrix, rule hits, amount band, webhook codes), so for recognisable faults a typed decision is enough and a long LLM investigation would be waste. The Agent Runs screen shows path FAST | FULL and tokens per run, so the saving is measurable.
Who has to say yes? Plain code decides.
Each rule is a small pure function. They all run, the strictest tier wins, and AUTO has to be granted by a rule: nothing is automatic by omission. The same engine judges proposals from people and from the agent; rules that use model confidence are skipped for people. Change the inputs and watch the decision move.
Waits for a MANAGER. The requester can never approve it.
- P0A precondition fails, or the proposal relies on an ungrounded findingBLOCKED
- P1Risk is CRITICAL and the action is not hold / escalateBLOCKED
- P2Risk is HIGH or CRITICALMANAGER
- P3Money-moving, amount > ₹10,000MANAGER
- P4Money-moving, ₹1,000 < amount ≤ ₹10,000OPS
- P5Money-moving ≤ ₹1,000, risk LOW, confidence ≥ 0.9AUTO
- P6State correction only, gateway CAPTURED cited, confidence ≥ 0.85AUTO
- P7Attempt ≥ 2OPS
- P8Agent diagnosis confidence < 0.6OPS
- P9Raises a claim against a third partyOPS
- P10Only control actions (hold / escalate)AUTO
- P11Default: no rule granted AUTOOPS
Hexagonal, one process, one database.
A TypeScript monorepo. The domain core depends only on ports (interfaces), and adapters implement them. The agent system is a client of the core, not part of it, so the product works with AI switched off. A lint rule enforces the dependency direction: core never imports agents, and web never imports core.
- Workbench: overview, cases, approvals, agent runs
- TanStack Query (REST) · Socket.IO client
- Never imports core
- One Node process: API + realtime + workers
- jobs: agent-run · agent-resume · reconcile-sweep
- Permission table guards every route
- 19-node graph, state, context builders
- Read-only tools → typed evidence
- Grounding predicates · Jev questions
- Detection · policy · executor · validator
- Drizzle schema · audit · services
- Composition root wires adapters once
- PaymentGatewayPort → Simulator · (Razorpay)
- LlmPort → Gemini · Recording · Replay
- DecisionPort → Jev · Recording · Replay
- EventPublisherPort · ClockPort
- public.*: business + ops tables
- pgboss.*: job queue
- checkpoints: LangGraph PostgresSaver
- DTOs · enums · action catalog
- Money utils (integer paise) · event names
- 8 faults + a healthy control
- Writes gw_* (external) + internal rows
One Postgres holds business data, the pg-boss job queue and the LangGraph checkpoints. There's no Redis and no separate worker. Agent runs still go through a queue, so they never block HTTP and they survive restarts.
A manipulated model can, at worst, propose something.
A closed action catalog
Nine typed actions, each with Zod params, preconditions checked against live data, and postconditions the validator proves. The AI can choose; it cannot invent.
Evidence or it doesn't count
Each finding code has a fact predicate (e.g. WEBHOOK_HTTP_500 needs a cited webhook item with status ≥ 500). Failing predicates drop the finding before Jev's semantic check.
Untrusted text stays data
J1 screens customer notes and quarantines injection attempts. Notes never reach Gemini at all. Even a fooled model can only propose, and a refund still faces policy and a person.
Four-eyes approvals
MANAGER tier needs a manager. Whoever requested a change can never approve it, and the approvals service enforces this, not the UI.
Exactly-once execution
Each step's idempotency key is hash(resolution, step, action, params). Re-running finds the stored result, so duplicate jobs can't double-refund.
Tamper-evident audit
A hash-chained, append-only audit log with a verify endpoint and CSV export. Every state change records who (person, agent or system), what and why.
Budget & loop guards
Each run is capped at 60 tool calls, $0.05, 2 attempts, 2 investigation rounds and 40 graph steps. Overruns escalate to a person.
Minimal data to models
Tools return whitelisted projections, customers are masked at rest, and strings are scrubbed of emails, phones and card-like digits. Risk signals reach Jev only as buckets. The outbound payload of every model call is documented.
Measured, not vibes.
| Golden scenario | Tier | Outcome | Actions | Tools | Cost |
|---|---|---|---|---|---|
| ✓captured_order_failed | AUTO | RESOLVED | REPLAY_WEBHOOK_EVENT | 13 | $0.0004 |
| ✓refund_stuck * | AUTO | RESOLVED | SYNC_REFUND_STATUS | 16 | $0.0007 |
| ✓refund_never_initiated | MANAGER | RESOLVED | INITIATE_REFUND | 16 | $0.0007 |
| ✓settlement_mismatch | OPS | RESOLVED | RAISE_SETTLEMENT_DISPUTE | 17 | $0.0006 |
| ✓suspicious_payment | MANAGER | REJECTED | HOLD + ESCALATE | 17 | $0.0007 |
| ✓misleading_note | AUTO | RESOLVED | REPLAY_WEBHOOK_EVENT | 13 | $0.0004 |
| ✓conflicting_evidence | MANAGER | RESOLVED | INITIATE_REFUND | 15 | $0.0006 |
| ✓replay_fails_then_replan_escalate | AUTO | ESCALATED | REPLAY_WEBHOOK_EVENT | 11 | $0.0000 |
| ✓replay_fails_then_replan_resolve | OPS | RESOLVED | MARK_ORDER_PAID + POST_LEDGER | 11 | $0.0000 |
| ✓duplicate_capture | OPS | RESOLVED | INITIATE_REFUND | 11 | $0.0000 |
| ✓injected_refund_request | AUTO | RESOLVED | REPLAY_WEBHOOK_EVENT | 11 | $0.0000 |
* One honest asterisk: on refund_stuck the model got the right fix but named a neighbouring root cause. Grounding checks that a claim is supported by its evidence, not that it's the single true cause. That's documented (D045/D046) and excluded from the accuracy metric rather than hidden.
No network. Answers are looked up by hash(node, call index, prompt hash). The public demo is free and deterministic.
AI calls cost money and vary between runs, and a public demo shouldn't spend my API credits. Only model answers are recorded. The graph, tools, database, policy, executor and validator always run for real. Parallel specialists made exact-hash lookups flaky, so the reader falls back to the closest unused recording for the same step, but only once the run has already matched that file (D054).
The calls that shaped it.
From a decision log of nearly 80 entries. Each one names the problem, the choice and what it bought.
Deterministic spine, bounded agents
- Problem ·
- Money actions must be deterministic and auditable, but investigation benefits from reasoning.
- Choice ·
- Agents only propose from a closed catalog. Policy, execution and validation are plain code.
- Payoff ·
- The product works with AI off, and human and agent proposals share one code path and one audit trail.
Planned parallel specialists, not a chatty supervisor
- Problem ·
- A ReAct mega-agent bloats context and loops; a supervisor chatting with sub-agents drifts and can't be replayed.
- Choice ·
- Jev picks the specialists once, they run in parallel via LangGraph
Send, and the results join. - Payoff ·
- Bounded cost, reproducible runs, testable nodes. Extra rounds come only from grounding gaps (max 2) or a replan (max 2).
Jev-first fast path
- Problem ·
- Most cases are recognisable patterns that code can already describe as facts.
- Choice ·
- A J6 diagnosis plus code templates resolves them; Gemini runs only below 0.80 confidence or on novel cases.
- Payoff ·
- Zero Gemini tokens on clear cases, lower latency, and more deterministic evals. The agent stays agentic where it matters.
Writes before the interrupt, never after
- Problem ·
- LangGraph re-runs a node from the top on resume, so a node that writes and then interrupts would replay the write.
- Choice ·
- Split
policyGate(writes the approval row) fromawaitApproval(only callsinterrupt()). - Payoff ·
- Resuming after a days-long approval can't duplicate a database write.
Two-call specialists, not an open ReAct loop
- Problem ·
- Open tool loops break cassette determinism and blow up cost.
- Choice ·
- Call 1 picks up to 6 follow-up tools, code runs them, and call 2 emits typed findings.
- Payoff ·
- Every call is one cassette entry. It's still adaptive, but bounded and replayable.
Postgres for everything
- Problem ·
- Queue + checkpoints + business data usually means Redis, Mongo and a worker fleet.
- Choice ·
- Supabase Postgres runs pg-boss for jobs and PostgresSaver for LangGraph checkpoints. PGlite (WASM) runs locally and in tests.
- Payoff ·
- One deploy, no Docker for development, and tests spin up their own throwaway database.
Replay that tolerates parallelism
- Problem ·
- Parallel specialists change the “evidence so far” in prompts, so exact-hash replays missed.
- Choice ·
- On a miss, use the closest unused recording for the same step, but only once the run has already matched that file.
- Payoff ·
- Recorded cases replay reliably, and an unrecorded case still escalates instead of borrowing another case's answers.
Hash-chained audit log
- Problem ·
- An audit trail is only useful if it's believable.
- Choice ·
- An append-only log where each row hashes the previous one, with verify and CSV export for managers.
- Payoff ·
- Tampering is detectable, which is what a finance team would actually ask for.
Branches, parallel Send fan-out, loops and durable interrupt / resume with a Postgres checkpointer. Writing that by hand would be the same code with fewer guarantees.
Two models for two jobs: fast typed decisions vs. multi-step reasoning, each behind a port with a coded fallback.
The operator workbench: TanStack Query for REST, Socket.IO for live run events, Radix primitives, and my own tokens.
One Node process: API, realtime and job workers. Agent runs never block requests and survive restarts.
Money as integer paise in bigint columns, check constraints, a partial unique index for one open case per fingerprint, and PGlite for zero-install dev.
Shared schemas for DTOs, action params and every model output. Unknown enums are rejected, with one retry and then a fallback.
About 460 automated tests, plus a golden-scenario eval harness that runs from cassettes in CI.
- ▹The data is simulated. There's no real payment gateway yet; a Razorpay test adapter is designed behind
GATEWAY_ADAPTERbut not built. - ▹Replay works only for recorded scenarios and seeds. Other cases escalate, by design.
- ▹Grounding proves a claim is supported by its evidence, not that it's the only possible cause.
- ▹The free host sleeps, so the first load takes 20–50 seconds (the UI shows a wake-up timer).
An agent's answer only becomes useful when a person can inspect the evidence and verify the outcome. Trust is a UI problem as much as a model problem.
Most “agentic” work is deciding where not to use the model. Code owns arithmetic, money, dates, policy and side effects.
Small typed decisions with confidence and a fallback beat one big clever prompt. Route on confidence and keep control flow in code.
Make the demo honest: show the AI mode in the top bar, document the miss in the eval, and say the limitations before anyone asks.