Playbooks, benchmarks, and results from the queue.
How teams put AI agents on their highest-volume work without giving up control. Grounded in real deployments — with the numbers, the flows, and widgets you can play with.
How Cenoa cut business onboarding from 2 weeks to 2 days
A digital bank for businesses turned a 2-week KYB backlog into a 2-day, day-one-revenue flow — with a compliance reviewer on every approval. The numbers, the flow, and what carried over.
How Getir ran 200K product onboardings a month with 5 people, not 25
6,000 images a day, a 25-person QA bottleneck, and GMV lost every hour a restaurant stayed offline. Here is the catalog-review flow that took time-per-task from 90 seconds to 1.5.
The cost of slow onboarding: 70% of firms lose a client over it
70% of financial institutions have lost a client because business onboarding was too slow (Fenergo, 2024). We break down where the days go and what each one costs.
SAP order management: from 15 minutes to 30 seconds per order
B2B order entry is copy-paste between email, PDF, and SAP. Here is how an agent handles the routine orders end-to-end and hands a human the exceptions, cutting 15 minutes to 30 seconds.
25 days to first sale: the marketplace activation benchmark
Across 65,000 sellers, average time from signup to first sale on enterprise marketplaces is 25 days (Mirakl, 2023). Catalog review is the hidden tax. Here is the math.
Intelligent document processing in logistics: same-day, 98% accuracy
BOLs, customs forms, and invoices arrive as PDFs and photos. Here is the document-processing flow that hit same-day turnaround at 98% accuracy — with humans on the exceptions.
$250 per exception PO: the real cost of invoice mismatches
A single purchase order that doesn't match its invoice costs $250 fully loaded to resolve (APQC). Multiply by your exception rate. Here's how to bring it down.
Getting past the 78% wall: the accuracy curve nobody shows you
Most agent pilots flat-line near 78% accuracy. The fix isn't a bigger model — it's supervisor corrections turned into training signal. Watch the curve climb, week by week.
Build in-house vs platform: 6–12 months vs 3 weeks
The build-vs-buy spreadsheet always undercounts the same things: eval infra, audit trails, and the engineering ticket behind every workflow change. A grounded comparison.
Clearing peak-season support volume without hiring for it
A flash sale triples the queue overnight. WISMO, refunds, and fraud holds pile up faster than temps can ramp. Here's the human-in-the-loop pattern that absorbs the spike.
It's not an automation problem. It's a deployment problem.
You bought a workflow tool. Your team configured it for months. The AI made errors nobody could explain, and the project got shelved. That's a deployment problem — here's the fix.
Catalog onboarding at scale: review thousands of SKUs a day
Every new SKU needs classification, copyright checks, and quality review before it can sell. Here's how to run catalog onboarding at scale with a reviewer on the edge cases.
An operating model, not a tool you buy
Workflow builders need rules. Internal tools need a developer. RPA breaks in production. Full agents leave no one accountable. Here's the operating model that avoids all four traps.
KYC/KYB onboarding without a compliance backlog
Document checks, sanctions screening, and risk scoring are the reason activation stalls. Here's how to clear the routine cases and put a reviewer on every flag.
The middle path: why augmented ops is the hardest thing to build
100% manual doesn't scale and burns out operators. Full automation is easy to break and impossible to govern. The middle — AI runs the flow, a human owns the last call — is the hard part.
Transaction dispute resolution at scale, with an audit trail
Chargebacks and disputes are high-volume, high-stakes, and deadline-driven. Here's how agents assemble the evidence pack and route the judgment calls to a human in seconds.
How the same team processes 10x the volume
The 10x isn't a bigger headcount or a faster tool — it's a flow where the agent moves optimistically through every case and the operator approves in one click. Here's how it works.
Shipment exception triage: catching SLA risk before the penalty
The 10% of shipments that go sideways eat most of the ops day. Here's how to triage customs holds, delays, and damage claims across the TMS, email, and carrier portals.
Why human-in-the-loop AI agents win in production
Autonomous agents stall at the trust bar. Human-in-the-loop ships. Here is the architecture that gets AI past the pilot stage and into real ops.
No rip-and-replace: put agents where the work already lives
SAP, Slack, HubSpot, your warehouse — the systems aren't the problem, the handoffs between them are. Here's how to deploy agents into your existing stack without a migration.
Intelligent document processing for freight paperwork
BOLs, customs declarations, proofs of delivery — the paperwork is the job. Here's how to extract, validate, and reconcile it with confidence routing and a clean audit trail.
What “production-grade” actually means for AI agents
A working definition of production-grade for AI agents: confidence thresholds, fallbacks, rollback, audit, and an eval harness that survives a model swap.
Supplier onboarding: from weeks to days before the first PO
Certifications, document checks, and compliance review across disconnected systems keep a new vendor waiting weeks. Here's how to compress that to days without dropping controls.
The eval harness gap: why agent pilots stall at 78% accuracy
Most AI pilots flat-line around 78% accuracy. The bottleneck is rarely the model. It is the missing eval harness. Here is how to build one.
Quality inspection triage before defects ship downstream
Inspection reports, defect photos, and return-material requests outpace the QA team. Here's how to triage them fast and put an engineer on the critical calls.
How to scope a 30-day AI agent pilot for ops teams
A practical 30-day playbook to take an ops workflow from messy spreadsheet to production agent without lighting compliance on fire.
Real-time payout and dispute ops for live events
When an event goes live, payouts, disputes, and support spike in minutes. Here's how to clear the routine cases in real time and route the risky ones to a human.
Confidence thresholds, explained: routing decisions to humans
Confidence thresholds are the steering wheel of a human-in-the-loop agent. Here is how to pick them, tune them, and avoid the common traps.
Content moderation at scale, with receipts
Moderation and creator disputes need speed and a defensible record. Here's how agents handle the clear cases in real time and hand humans the judgment calls, logged.
Audit trails for AI agents: a compliance primer
What an auditable AI agent actually logs, why it matters for SOX, GDPR, and HIPAA reviewers, and what to ask your vendor before you sign.
First-pass claims adjudication in minutes, not days
FNOL, document intake, and policy checks are repetitive until they aren't. Here's how to adjudicate the clean claims end-to-end and hand adjusters the ambiguous ones, evidence-ready.
Why supervisor corrections are the moat
Compounding accuracy is the only loop we have seen consistently beat the 78% wall. The fuel is supervisor corrections, structured into training signal.
Choosing your first ops workflow for AI automation
A scoring rubric to pick the workflow with the best chance of clearing a 30-day pilot, without betting the quarter on a moonshot.
Tier-1 ticket triage with HITL agents: what to expect
Volume, accuracy, escalation rates, supervisor load. A real-world look at running tier-1 ticket triage on a human-in-the-loop agent.
Vendor onboarding automation: where the bottlenecks actually live
Most vendor onboarding teams blame procurement. The real time-sinks are paperwork parsing, sanctions screening, and ERP data entry. Here is the fix.
First-pass claims review: patterns, anti-patterns, and pitfalls
What separates a claims review agent that actually ships from one that lives forever in a sandbox. Patterns, anti-patterns, and the pitfalls we have hit.
Lead enrichment with AI: scores you can defend in QBR
An ICP score nobody trusts is worse than no score. Here is how to build lead enrichment that reps lean on and that survives a quarterly review.
Self-built vs platform: the real cost of in-house agent infra
We built it twice. Here is what we underestimated, what surprised us, and the question to ask before you spin up the build-vs-buy spreadsheet.
Per-seat vs volume pricing for AI agents (and why we chose volume)
Per-seat pricing punishes the workflows AI agents are best at: high volume, low headcount. Here is the math behind volume-based pricing.
When NOT to use an AI agent: a checklist for ops leaders
Not every workflow is an agent workflow. A six-question checklist to help ops leaders decide when to ship a HITL agent and when to walk away.