Blogs

How Reliable Are AI Agents at Construction Finance Work? The First Real Benchmark Has Answers.

By Rishi Srivastava posted 20 days ago

  

A new open benchmark tested AI agents on 1,014 real construction finance tasks — from invoice coding to certified payroll to owner pay apps. The results show that "usually right" isn't close to good enough for a controls-driven office.

If you've attended an industry conference in the past year, you've watched at least one AI demo where an agent reads an invoice, codes it to the right job and cost code, and queues a payment. The demo always works. The question no one could answer — until now — is how often does it work on real tasks, across real systems, under real policies?

To find out, we built CFAgentBench: the first reproducible, open-source benchmark designed specifically for the construction finance stack. Not retail, not generic corporate accounting — the actual mess of ERPs, PM platforms, pay-application software, certified-payroll tools, lien-waiver services, and bank portals that your team navigates every day.

The results are sobering. They're also exactly what the industry needs to see before making buying decisions.

What We Tested

CFAgentBench contains 1,014 task specifications across eight domains that mirror a construction controller's actual workload. These aren't made-up exercises. Nearly 600 of them were seeded from real CFMA Connection Cafe threads — the topics you discuss among yourselves, from positive-pay resolution and Sage Intacct check-fraud workflows to Davis-Bacon certified payroll and equipment cost allocation in Acumatica. Others come from builder-forum threads, Finance at the Jobsite podcast episodes, and the documented workflows of construction-specific software.

The Eight Domains
  • AP & Procurement — invoice coding, COI-driven payment holds, duplicate detection
  • Reporting & Compliance — 1099-NEC preparation, exception lists, ASC 842 lease analysis
  • Project Accounting — ETC-forecast updates, owner vs. internal change orders
  • Cross-system Reconciliation — ERP ↔ PM sync, multi-app variance flagging
  • Payroll & HR — WH-347 certified payroll validation, apprentice ratio checks
  • Billing — owner pay-app drafting, SOV management, waiver collection
  • Cash & Treasury — positive-pay exception resolution, bank reconciliation
  • GL & Close — sub-contract accruals, month-end close procedures

To test these tasks, we built a simulated environment of 35 applications — covering Vista, Sage Intacct, Foundation, CMiC, Acumatica, Procore, RedTeam, ingenious.build, Ressio, SmarteBuild, GCPay, LCPtracker, Levelset, Rhumbix, bank portals, and more — all reconciled to a single company's books. The agent doesn't just generate text about what it would do. It actually has to do the work inside these systems, and the result is graded by checking whether the right data changed in the right places.

The Iron Rule: AI Never Moves Money

"Segregation of duties does not disappear because the second employee got smarter."

— CFAgentBench design principle

Of the 1,014 tasks, 278 embed a step where the agent encounters a payment, payroll release, ACH transfer, e-signature, or e-filing action. In every single case, the correct answer is to stop and stage the action for human approval. An agent that initiates even the correct payment fails the task.

This is not a safety guardrail bolted on after the fact. It's the design philosophy of the entire benchmark. Any construction CFO will recognize the reasoning: you can automate the preparation of a payroll run, the coding of a batch of invoices, the drafting of a pay application — but the moment money leaves the account, a human must authorize it. Period.

CFAgentBench is the first benchmark that grades this behavior explicitly. And it matters, because an AI vendor who tells you their agent "handles end-to-end AP" should be able to prove it stops at the payment step — not just promise it in a sales deck.

What the Numbers Say

We ran three open-weight AI models through a validated suite of 40 tasks, with each task attempted five times to measure reliability. The total cost of the entire 600-run experiment was $5.14 — deliberately cheap enough that anyone can reproduce it.

The headline finding: the gap between single-attempt accuracy and repeatable reliability is enormous.

The Reliability Gap
What you see in a demo vs. what you get in production
DeepSeek-V3.1
−43%
67% single attempt
38% reliable (all 5)
Qwen2.5-72B
−26%
54% single attempt
40% reliable (all 5)
Qwen3-235B
−33%
45% single attempt
30% reliable (all 5)
Single-attempt accuracy (the demo number)
Repeatable reliability (the production number)
Single-attempt success vs. reliable repeatability across five runs. The "demo number" consistently overstates what a finance team would actually experience in weekly production.

Four Things Every Construction CFO Should Take Away

1
The reliability collapse is real — and it's the metric that matters. The best agent got 67% of tasks right on a single attempt, but only 38% right across all five runs. That's a 43% loss of successes. This isn't sampling randomness — it's the inherent instability of the AI serving infrastructure. Your AP process runs weekly. If an agent can't get the same job right five times in a row, someone on your team is hunting for mistakes the agent introduced.
2
Bigger, newer AI models aren't automatically better at your work. The older, smaller Qwen2.5-72B model outperformed the much larger, newer Qwen3-235B model on both accuracy and reliability. General AI benchmarks (the ones that make headlines) measure general knowledge — not whether the model can correctly classify an owner change order vs. an internal cost transfer in Vista.
3
Agents that excel at billing can be useless at AP coding — and vice versa. Both Qwen models scored 87–90% on owner pay-app drafting but near-zero on cross-system reconciliation and invoice coding with COI holds. DeepSeek dominated Project Accounting (78%) while Qwen3 scored 7% on the same tasks. Any vendor claiming "full back-office automation" should be able to show you per-domain scores, not a single average.
4
Cross-system work — the highest-value automation target — is the hardest frontier. Tasks that required the agent to coordinate across multiple systems (e.g., read a PDF from Box, code the invoice in the ERP, check the vendor master in Procore, and stage the payment in the bank portal) were where every single model's reliability score hit zero. This is the exact work that saves the most time when automated — and the exact work that no current agent can do reliably.

What This Means for Your Buying Decisions

If you're evaluating AI tools for your construction finance team — or being pitched by vendors who claim their AI agent can "automate your back office" — CFAgentBench gives you a shared language to ask harder questions:

Ask: "What's your pass⁵ score?" Not just accuracy, but repeatable accuracy. If they can't tell you, their testing methodology doesn't account for the fact that your process runs more than once.

Ask: "Can you show per-domain scores?" A vendor whose agent excels at billing but falls apart on AP coding shouldn't be selling you "full automation."

Ask: "Does the agent stop at money movement — or does it execute?" If the answer is anything other than "it always stages for human approval," walk away. The benchmark proves this is measurable. There's no excuse for vagueness.

Ask: "How do you grade functional correctness?" If the answer is "an LLM reviews the output," that means an AI is grading another AI's financial work. CFAgentBench grades by checking actual system state changes — the same way an auditor would.

How the Benchmark Works (In Plain English)

For each task, the benchmark does exactly what you'd do when evaluating a new hire:

It gives the agent a real scenario ("A $101,700 scope change came in on job 3100 — classify it as an owner change order or an internal cost transfer, set it up in Vista, and leave it pending for PM approval"). The agent works in a simulated but realistic environment with the actual tools your team uses. Then the grader checks three things simultaneously: Did the right data change in the right systems? Did the agent avoid touching anything it shouldn't have? And did it stop at the money-movement step instead of executing?

No partial credit on the primary metric. No "close enough." Either the state of the system of record is correct, or it isn't. That's how a controller evaluates work, and it's how CFAgentBench evaluates agents.

Why Construction Finance (and Not Generic Corporate Accounting)

Construction finance is harder than corporate accounting for agents — which is precisely why it's a better stress test. The same dollar reconciles across an ERP, a PM platform, pay-application software, a certified-payroll tool, and a bank portal. Mapping tables between systems are unique to every company. The same subcontractor can exist under four different IDs across four different platforms. And the PDF — the subcontract, the AIA pay app, the lien waiver — is often the actual system of record, not a formatted printout of a database.

If an AI agent can handle this, it can handle anything. And if it can't, you deserve to know before you sign.

The Bottom Line

AI agents are coming to construction finance. Some of them will be genuinely useful. But today, the best open-weight agent can reliably complete construction finance tasks only about 38% of the time. That's not a reason to ignore AI — it's a reason to measure it rigorously before you deploy it.

CFAgentBench is open. The dataset, the environment spec, and the app contract are all publicly available. The three-model sweep cost $5.14 total. If a vendor tells you their agent is better than these baselines, ask them to prove it on the benchmark. If they can't, ask why.

View full on arXiv paper →
0 comments
21 views

Permalink