Skip to content

AI systems

Razor: risk layer for AI-agent payments

Catches what fraud tools miss: an AI agent with a valid payment mandate that was hijacked into buying the wrong thing.

Role
Personal project
Type
AI systems
layers: mandate, behaviour, intent, evidence
4
ROC-AUC, held-out synthetic data
0.87
attack classes simulated
6

How it fits together

  1. Agent sessionCart, delegated mandate and prior history
  2. verify
  3. L1 MandateAmount, category, time, lifecycle and identity checks with a JSON reason trail
  4. score
  5. L2 BehaviourSequence, timing, token reuse and velocity, scored with a risk model
  6. check purpose
  7. L3 IntentPurpose against cart, injection patterns, beneficiary novelty
  8. escalate
  9. L4 EvidenceLiability call, evidence packet, grounded narrative, step-up or mandate pause

The problem

Agentic checkout protocols deliberately leave merchant fraud modelling out of scope. Legacy signals such as device, IP and clickstream describe the agent, not the person behind it, so honest automation can look like a bot while a hijacked agent can still make a purchase that sits inside its mandate. The new failure is indirect prompt injection: an agent reads attacker-controlled content and then buys something unrelated to its task.

How I approached it

  1. 1

    Was it authorised?

    L1 verifies the delegated mandate (amount, category, time window, lifecycle, identity) and returns pass or fail with a JSON reason trail. Razorpay has no first-class AI-agent object, so delegation is modelled as a mandate: a scoped, revocable, user-approved grant.

  2. 2

    Did the session behave normally?

    L2 scores sequence, timing, token reuse and velocity, and returns a risk score with the features that drove it.

  3. 3

    Did the outcome match the purpose?

    L3 looks for divergence between the stated purpose and the cart, injection structures, and beneficiary novelty combined with high value and timing escalation.

  4. 4

    What is the evidence?

    L4 makes a deterministic liability call and builds an evidence packet with a grounded narrative, then triggers a step-up, a mandate pause or the Razorpay dispute flow.

What I built

  • The headline demo (A6) is a valid, in-limit mandate on a normal-looking session whose cart (a crypto voucher, a luxury watch, a gaming console) no longer matches a grocery purpose. L1 and L2 pass, L3 flags it, and L4 escalates to the provider.
  • A simulator with six attack classes, from A1 consent replay to A6 injected intent, and a browser console that shows each case layer by layer.

The result

On a frozen, mandate-disjoint held-out split, Layer 2 reaches ROC-AUC 0.867 and PR-AUC 0.611, and catches 100% of spoofed-identity, over-ceiling and slow-drain attacks. It deliberately scores injected intent (A6) at 0%, because those sessions look normal; Layer 3 handles that class at 100% precision and 65% recall. All figures are on synthetic data, not production fraud rates.

Built with

  • Python
  • FastAPI
  • Razorpay
  • scikit-learn
  • XGBoost
  • LightGBM
  • SHAP
  • Pydantic
  • pytest