AI Agents in Construction Finance: What Actually Works (and Where They Break)
Summary
Rishi Srivastava opens with a promise: no product pitch, just an honest look at where AI agents genuinely help construction back offices and where they fall apart. Speaking to a room of controllers, CFOs, specialty contractors, and ERP partners, he frames the entire hour around trust, controls, and audit survival. Agents work well on bounded, reviewable tasks like reading messy pay apps, coding invoices, and migrating data — but they break on anything with a deadline, ambiguity, or system glitches that require unattended recovery.
Grounding the discussion in research, Rishi walks through Meta's GAIA-2 benchmark, where the best frontier model completes only forty-two percent of realistic everyday tasks autonomously. He then introduces CF Agent Bench, a benchmark his team just published that tests over a thousand construction finance tasks across thirty-five mock apps. The most sobering finding: agents that succeed two-thirds of the time on a single run drop to thirty-eight percent when asked to repeat the same task five times in a row — exactly the reliability profile weekly workflows demand.
The episode explores how the Model Context Protocol (MCP) is dismantling vendor lock-in by making data portable across ERPs and AI tools. Rishi demonstrates an agent migrating data from Sage 50 and another creating a Vista job from an award email, showing how the plan-act-observe loop keeps humans in control of every write. He argues the real switching cost was never the software — it was the trapped data — and open protocols now let finance teams own the control layer while keeping a swappable stack.
The final stretch turns into an open conversation with attendees. Dwight, Matt, Veronica, Doug, Jayna, and others share war stories and adoption philosophies, from starting with the lowest-risk internal Q&A tasks to building checksums that guard against hallucinated data. The consensus: trust has to be earned transaction by transaction, humans stay on anything money-moving or deadline-driven, and the smart path forward is bounded workflows, narrow MCP data slices, and proving accuracy on your own documents before scaling.
Key moments:
- The 42% Reality Check on AI Agents: Rishi reveals that even the best frontier model completes only forty-two percent of realistic everyday tasks autonomously on Meta's GAIA-2 benchmark, reframing expectations for construction finance leaders.
- Where Agents Break: Deadlines and Ambiguity: Time-sensitive tasks score near zero because agents think too slowly to beat the clock, making pay app cutoffs, lien waiver windows, and draw schedules dangerous for unattended automation.
- CF Agent Bench Reveals Reliability Collapse: Beiing Human's newly published benchmark shows agent accuracy drops from sixty-seven percent on single runs to thirty-eight percent when repeating the same task five times in a row.
- MCP Ends Vendor Lock-In: The Model Context Protocol acts like USB-C for AI, letting finance teams move data between ERPs on open rails and own the control layer instead of being trapped by proprietary systems.
- Live ERP Migration and Job Creation Demos: Rishi shows an agent extracting Sage 50 data for migration and another drafting a Vista job record from an award email, with humans approving every write to the system of record.
- Money Movement Guardrails as Scored Outcomes: CF Agent Bench includes 278 tasks where the correct answer is to stop and stage for human review, making segregation of duties an auditable outcome rather than a hope.
Listen on Spotify, Apple Podcasts & Audible
Transcript
[00:00:00] Welcome to Finance at the Jobsite, the podcast where construction finance meets the field. I’m your host, Rishi Srivastava, founder of Beiing Human. In each episode, I sit down with construction CFOs, controllers, owners, project managers, IT leaders, ERP consultants, and industry experts to uncover how they connect the back office with the field, choose and implement technology, manage cash flow, and drive profitable projects.
Whether you are running the numbers, leading the team, or designing the systems that keep projects moving, this is your place to learn what’s working, what’s broken, and what’s next in construction finance okay, we’ll begin Thanks for giving me an hour. Uh, quick promise up front, this isn’t a product pitch dressed up as a webinar. I run a construction finance software company, so yes, I have skin in the game. But what I wanna s– do today is give you the honest version of where [00:01:00] AI agents actually help in a back office and where they fall on their face.
I use these tools every single day. Claude, a couple of agents I built, ERP migration tooling. So most of what you’ll see is stuff I’ve actually run, not slideware. Uh, and I’ll back it with two pieces of research, , Meta’s big public benchmark, and a paper my team just published on agents doing CFO-grade work.
Let me tell you who’s actually here, uh, because it shapes how I’m going to talk. This isn’t a general tech crowd. The bulk of you sign the checks and own the close. Controllers, CFOs, folks at, uh, specialty contractors and GCs in that ten to a thousand employee range. specialty subs here, electrical, mechanical, concrete, marine, plus some – CFMA [00:02:00] leadership and ERP partners in the mix.
So I’m going to not wow you with novelty. The currency in this room is trust, controls, and whether something survives an audit. That’s the lens I’ll use the whole hour, and I’ll then tell you agents work. I’ll show you an agent doing the ERP migration and an agent creating a job from an email.
So here’s the shape of the next hour. Uh, we start with workflows directly applicable to you, then get specific. What genuinely works right now, and then the part nobody likes to put on a slide where these things still break. We’ll ground it in the research and bring it home to your world, ERP, AP, project management, and finish on where this is headed in the next year or two, and how you get your team in front [00:03:00] of it.
Three polls along the way to keep me honest and keep you awake.
So, uh, the hype says agents are about to run your whole back office unattended. That’s not what the research says, and it’s not what I see in my own day. The truth is more useful and more specific than the hype, and frankly, more useful than the skeptics who say it’s all nonsense. On the left, uh, the stuff that already works.
Think reading a messy scanned pay app and pulling the right numbers out, coding an invoice and matching it to the right PO and job, moving data from one system to another, drafting and summarizing. Notice the pattern, bounded, repeatable, and a human can check the work
On the right, uh, where it breaks, anything on a clock, anything where the situation changes halfway through, anything ambiguous, anything where the system glitches [00:04:00] and goes quiet, and above all, anything unattended and irreversible. The whole game today is putting agents on the left column with the person owning the right.
That’s not a limitation to apologize for. That’s actually the design
Okay, time to wake up. Uh, okay, first poll. The single best AI agent in the world right now on realistic everyday tasks start to finish, no human help. What share does it actually complete? Ninety percent, seventy-five, forty-two, or twenty-five? Drop A, B, C, or D in the chat. Be honest, not optimistic So I’m gonna watch what’s going on in the poll
So let’s see. C, D, C. Seeing lots of Cs and Bs. Okay. Anyone else wants to put in three C? See lots of [00:05:00] Cs. Uh, we got some really smart people here in this, I feel like. So let’s go back to the deck one more time
So the answer is C, about forty-two percent. That’s the best model, GPT 5 as of last fall. Claude was around — thirty-five. Uh, the best open source options around twenty-one. So the frontier of fully autonomous everyday task completion is a coin flip. Worse than, a coin flip, actually. This is from a benchmark called GAIA-2, Meta’s benchmark from last year.
And the reason I trust it, it’s not trivia. It’s around a thousand human-written scenarios in a simulated phone, email, calendar, contacts, files, about a hundred tools. The agent has to read and write, take real [00:06:00] actions. And here’s the cruel part. The world runs on its own clock. Events fire , whether the agent is ready or not, just like your month-end.
A competent human scores basically a hundred. No special knowledge needed. It’s purely a test of reliability under real conditions, and the best machine gets forty-two. So let’s talk about the shape of the failures. Now, where it fails is the useful part. Break the tasks into categories, and a clear picture emerges. Following explicit steps, basically solved. Searching for information, basically solved. But the moment things get ambiguous or conditions change mid-task or a tool glitches, performance drops hard
And the worst category by far is anything time-sensitive. On the hardest deadline tasks, even the top models score near zero. Not [00:07:00] because they’re dumb, because they think too slowly to act before the deadline passes. Here’s the tell. Take away the thinking time penalty, and the score jumps to around thirty percent Read that, back into your world.
The agent isn’t bad at the pay app or invoice or payroll. It’s bad at beating the clock, the cutoff time
Here’s the part that surprises people, and, and it matters for your budget. The instinct is, if it’s not reliable enough, throw a bigger or smarter, more expensive model at it. The data says that that flattens out fast. More reasoning, more compute, diminishing returns almost immediately. And worse, the smarter models are often slower, which on a deadline task actually hurts you So, uh, stop shopping for the highest score on some leaden leaderboard.
Uh, [00:08:00] shop for correct per dollar, the cheapest, fastest model that clears your accuracy bar. And critically, prove it clears that bar on your documents, not on a benchmark. That single reframe will save you real money
Let me pull back the curtain on what an agent even is because the world’s overloaded It does not answer in one shot. It loops the exact way a good ERP migration consultant works. They don’t guess the whole project. They check the actual data at every step. It plans, figures out QuickBooks data map into Vista.
It, acts… actually moves a chunk. It observes, do the trial balance tie out, then it feeds that back into the next plan. Again, plan, act, observe, feedback. Around and [00:09:00] around
Three reasons, uh, this loop is built for your world. Mm. It checks real data at each step, so it hallucinates less. And every step is logged, so it’s auditable, and it stops when something’s off instead of barreling ahead. That loop is the whole reason any of this is actually trustworthy
Now again, have a poll to keep you guys, uh, awake. Now that you’ve seen the shape of it, which task makes today’s agents fail the most? Pulling data from a clean PDF, searching the web, acting before a deadline, or following explicit steps?
Vote in the chat, and we’ll see if the research agrees with you.
It’d be lots of, lots of Cs Nice. Okay, let’s go back
M-most of you got it. It is, it’s C. Uh, so I’m gonna come back to the [00:10:00] answer. Again, it’s C, so anything on a clock. But let me make all of these real because they are not abstract for you. Uh, go one by one. Um, so deadlines, uh, like pay app cut-offs, lien waiver, windows, draw schedules. You know, moving targets. A change order lands while the agents make tasks.
PO gets revised, scope shifted. It doesn’t handle the rug pull well. Ambiguity. Two vendors with nearly the same name, a missing PO, coding rule thats all exceptions. M-mess confuses it. Flaky systems. The ERP API times out, a sync is half done, a file’s malformed. It struggles to recover gracefully. This is exactly why human-in-the-loop isn’t a weakness, it’s the architecture.
The agent does the grind, the person [00:11:00] owns the judgment and the clock
So here’s the real one, uh, from my week. We win a job, let’s say all the details, number, owner, contract value, dates, sc-scope, they’re sitting in an award email. Normally, somebody hand keys that into Vista. Instead, I just tell the agent, “We won this. Make the Vista job from my email.” It reads the email, pulls the fields, drafts the job record, and then it stops.
It shows me every field. I review. Only on my say so does it actually create. Again, the loop we just talked about, layout, plan, act, observe, feedback, and nothing is written until I approve. Again, that approval gate is the whole point
Uh
I’m going to play a short clip, uh, uh, VC’s take because it reframes why this matters strategically beyond productivity. The [00:12:00] one-line version before I hit play, when an agent can rebuild your entire ERP setup in a day, the switching costs that locks you into software just evaporate.
Dude, where the fuck is value in a world of Anthropic and Claude Code wiping billions of dollars off a stock market? How should I think about that? Uh, you should think that software, cost of ca- creating software is going down to zero. That’s it. So, uh, and that means that, like, everyone will be able to generate software at any point in time.
So it is a massive change, and, uh, I was a hundred percent con, you know, convicted about this already when I saw this one or two years ago. So I’ve been– That’s been very, very clear to me. If the cost of software creation is going down, how do we determine which businesses have sustaining value versus which do not?
So the, the key thing right now is, so far, the only thing that’s gone down to, or not to zero yet, but become extremely [00:13:00] much cheaper, is the generation of software. The next thing that’s gonna hit everyone bad is the switching cost of data. Because so far, what you’re seeing is you have proprietary data stuck in, for example, the CRM vendor or the, other software as a service that you’re using currently.
So you may replicate and build the same dashboard or build the same processes in your own tool, but all your data is in there according to their data model, according to their setup. What’s gonna happen is people are gonna start solving that problem. How do I get all of my data from the existing vendor and move it to the new vendor with the help of AI through one click?
That brings down switching cost, and that’s when the real threat to SaaS comes. So let’s unpack that. Here’s the thing about the vendor lock-in. It was never really the app. You can always find another app. The lock-in was the data and your custom integrations trapped inside one vendor.
That’s the moat. So, um, [00:14:00] what’s changing? It is an open standard called MCP. Easiest way to think of it is USB-C for AI. You build a connector once, and it works across Claude, ChatGPT, Gemini, whatever. You’re not rebuilding for each one. Which means, uh, you can own the control layer and keep a swappable stack, change ERPs without burning a year.
And the proof that this is real, not theoretical, is that migration itself has become an agent task. So the question I would put to every vendor to deal with, including me: Can my data leave on open rails, or am I locked in? Make them answer it
So this is the workflow I’m most excited about, and it’s a textbook case of what agents are good at. Bounded and reviewable. Map the fields, move the history, reconcile, flag the exceptions that humans should look at. So for example, QuickBooks Desktop into Acumatica, QuickBooks in-into Vista, Sage 50 [00:15:00] into Acumatica.
They used to be a long hand-keyed ordeal that everybody dreaded. Now it’s an agent task, and one of them will be open-sourced as well. So you’re welcome to take a look at it. Uh, let me show you that one actually running
So extracting the data from Sage-50
And then start the migration. A-again, uh, in a real workflow, uh, there may be exceptions where a human will come in. Uh, but most of the dreaded, part data migration is automated
Okay. Now Uh, so GAIA two is a phone benchmark. Brilliant. But it’s emails and calendars. Nobody had built the equivalent for your books, so we did it, and we published it, uh, this month. It’s, uh, published. It’s live on, uh, archive. Um- CS.ai, uh, journal. Mm-hmm. So it’s called [00:16:00] CF Agent Bench. Over a thousand tasks across the real finance stack, ERP, project management, AP, payroll, certified payroll, lien waivers, bank portals, Surety.
And the key design choice is we grade on what actually changed in the system of record, not on what the agent claimed it did, but on what it actually did. So we didn’t make these tasks up. Uh, mm, they… most of them are sourced from, uh, CFMA connection, Cafe thread, or from the a podcast I host, from our actual customer conversations, emails, and built around the standards you live in, G seven oh to G seven oh three, W S three forty-seven, thousand ninety-nines, lease accounting.
And it run again, it runs against thirty-five mock apps, Vista, Intaact, Foundation, Procore, that reset every run, graded on whether the task functionally worked. [00:17:00] So and my favorite piece is the money movement guard. Two hundred and seventy-eight tasks where the correct answer is to stop and stage it for human review.
If the agent executes the payment even perfectly, it fails the task. So we made segregation of duties a scored outcome, not a hope. Just because, uh, one employee got really smart, it doesn’t mean that the segregation of duties disappear.
That’s the thing an auditor wants to hear.
So here’s the finding I tattoo on every implementation plan. Run a task once, the best model gets two-third of them right. Sounds great. Now make it do the same task right five times in a row, the way any weekly process has to work, it drops to thirty-eight percent. It loses almost half its win just by having to repeat them.
And that’s not one model, actually. Every agent we tested degrades the [00:18:00] same way from single shot to five in a row. The trivial baseline that does nothing scores zero as it should, so the benchmark is honest. So… And there is a hard frontier where it gets brutal. Cross-system sync, AP invoice coding with COI payment holds.
Every model drops to basically zero. These are exactly the high-value repeated workflows you’d most want to automate. So the lesson isn’t don’t automate, it’s demand the five in a row number, not the demo number. Oh, and the whole sweep costs just about five dollars to run. So there’s no excuse not to retest every time in a new model ships.
So we’ve intentionally not tested it on, uh, like really expensive models like, uh, Fable or Opus 4.8, so this can be a repeatable test. Uh-
Now a little bit, uh, about the model [00:19:00] control protocol. So one piece of, uh, how it works is it actually explains the accuracy numbers. An AI ha-has a fixed – attention budget. And think of it as a limited amount of focus. Dump your whole job cost table in front of it, and it loses the thread and gives you wrong numbers.
Hand it a narrow slice, and it’s fast, accurate, and you can audit it. That’s what MCP is. So the way you do this is basically a saved query, the AI calls by name, where you predefine exactly which fields come back so it’s physically can’t overpull. So ask for job costs on job one eighteen, and you get the five clean fields, basically: budget, actual, committed, not five thousand rows of ta- noise.
Constraining what it can see is what makes it reliable. So again, that, that tells you that less [00:20:00] is more, literally
Here’s all of that paying off and something you’d actually want Monday morning. An agent pulling from three different systems, uh, let’s say your ERP, your estimating, your field reports, and through those narrow MCP slices and producing a finished weekly job cost variance report at seven AM every M-Monday before anyone’s had coffee.
And it doesn’t just summarize actually. It flags tasks, order bleeding, uh, forty-two grand and tells you the fix, file the change order claims before the contract deadline. It surfaces unbilled work before it ages into a write-off. And you can just ask it to follow up in English. No dashboard, no pivot tables.
That’s the, that’s the works today bar. Bounded, reviewable, genuinely useful.
S-so this isn’t me predicting the future. It’s actually [00:21:00] already happening. In fact, very fast. Job trade, uh, sh-shipped a native AI connector in April this year. Twenty-five hundred companies connected in under two months, their fastest-growing integration ever. inGenious.global built first-party MCP across schedule, budget, RFIs, contracts, invoices, and importantly, it, it honors each user’s permission.
And as of, uh, this month, that includes us. We just launched our own MCP server so you can point Claude or ChatGPT or Gemini at Beiing Human with every change validated and logged. The pattern’s clear. Um, your project and finance systems are becoming places agents can act, not just read.
So here’s a workflow nobody loves, but everybody owns: permitting. It’s the perfect stress test for an agent because it’s everything we’ve been talking about in one place. It’s [00:22:00] repetitive and form heavy, which is the part an agent is great at. But it’s also on a hard clock and spread across a dozen jurisdiction portals that all work differently, which is exactly where they get dangerous So the right shape here is let the agent do the grind, gather the documents, prefill the application, track out expiration and renewal windows, flag what’s coming due.
But the actual submission, the thing with the deadline and the legal weight, a human presses the button. The agent makes sure you never miss a permit deadline. It doesn’t get it to file on its own. Same pattern as everything, mm, else, uh, I’ve been talking about. Agent on paperwork, human on the clock.
So n-now I know some of what, what some of you are thinking. Half my systems don’t have an API. This is irrelevant to me. Watch this. When there’s no API to connect to, the agent [00:23:00] just uses the computer the way a person does. Clicks, types, navigates the sc-screens. It’s a bridge for all the legacy no API tools construction still runs on.
Where the clean connection can’t reach, this can. I’m Puja, and I’m a researcher at Anthropic
I’m going to show you a simple example of computer use today. My friend’s coming to San Francisco next week, and I wanna take him to do some touristy stuff. I think doing a sunrise hike with a view of the Golden Gate Bridge never gets old. So I’ll ask Claude to figure out some logistics for us. I’ll ask Claude to find a good place to see the sunrise, to help me figure out timing logistics, and help drop a calendar invite so I remember when I have to leave.
It’s opening Chrome, going to Google, searching
It looks like it’s found something. Great. So how far away is the location from my place? It’s opening Maps
Searching for the distance between my area and the hiking location
Cool. So now it looks like Claude is searching for the sunrise time tomorrow
And is now dropping it into my calendar
and populating it with some [00:24:00] details
It looks like Claude did it. This is a simple example, but we’re sharing computer use early to learn from what people build
Quickly. So this is our, uh, third poll, and this, in my view, is, uh, hmm, the most important one. I would love, uh, people to participate in this one. So, uh This one’s a real conversation. Where would you let an agent run today? Pick all you’re comfortable with. Draft invoice coding for review, migrate history to a new ERP, post to the ERP unattended, or file something on a hard deadline alone
A, D. Matt saying A, D. Uh, Craig’s– Nate saying A. Steve saying A and B
A and B. Nobody said D
Actually, Matt said D, I think. Did anybody say D? M-Matt, uh, Matt, would you, uh, elaborate, uh, what you’re thinking? Um, [00:25:00] I, I think there’s a lot of, you know, deadlines with regards to, um, you know, fairly simple things, you know, COIs and things like that. Where they just need to get, uh, you know, sent to customers or vendors.
I think that’s a task that, it’s not a… I don’t, I don’t think– I think we could do it alone, you know, let the AI do it alone just because, um, I don’t think there’s a, a, a downside really. If it, if it misses, we just get reminded. Um, you know, so for that kind of a task, I think it would work, right? Ah, okay.
Okay. That makes sense. Are there some, uh, tasks in the category D that you would not let an AI run?
Um, I think something w-you know, more like, you know, my payroll taxes, um- Mm-hmm … or something of that nature where there’s some significant ramifications if something isn’t right. I think those kinds of things just need to be reviewed [00:26:00] or, even, payments like payroll tax payments, things of that nature.
I’d want to make sure that they’re verified before they get done right. Yeah. Yeah. Makes sense. So A and B definitely lit up, and C and D mostly they stayed in dark. So Mary, uh, do you wanna tell us what type of tasks you are thinking for C, Mary Devolt?
M-Mary or Alvaro a-also said D. I’m kind of interested because D is, uh, a task that’s, uh, kind of risky in some ways. So b-basically, if you’re looking at C, post to the ERP unattended, that’ll be- Mm-hmm … transactions that went across, um, different modules, automating the posting after it’s gone through. Let’s say you’re doing, um, uh, pro forma invoicing or whatnot, and that process- Mm-hmm
completed, having it automatically post to the AR, for example, would be something that wouldn’t be, in my opinion, uh, uh, a [00:27:00] problem. And who, who was speaking? That was- That was Mary- Mary, thank you. Yeah. Anybody else who wants to give us, uh, their take? Thank… Nate, what are you thinking?
I’m kind of picking on people, but I would love somebody giving some insights here
Hey, I’ll, uh, go ahead and give you kind of a rundown of what I was thinking. Mm-hmm. This is Veronica. So you can do just about anything with it, as long as you have, um, if it’s… Uh, obviously agents are created on a loop, so it’s going to keep going back and forth, back and forth until it’s correct, which is the difference between agents and, like, your chatbots and stuff like that.
So as long as you have that in, plugged in- Mm-hmm … um, then as long… It’s going to check to make sure that it’s correct before it does anything further. So that’s kind of a stopgap between, you know, everyone saying, [00:28:00] “I want to make sure that it’s correct before it posts to either the ERP or the invoice goes out,” or anything like that.
Um, that’s kind of my take on, you can do almost anything you want with it- Mm-hmm … and however you want it to post or, you know, leave it in the save but unposted so you can audit it later. The ERP has an audit trail so you can go back and say, “Yep, all of these things were done com- you know, correct.” And so that will kind of help you on that, that stopgap of, of the agent Yeah.
Yeah. Wherever you have a clear audit trail and, uh, the consequences maybe are mi- minimal for an agent to make mistake. Yeah. I don’t know how many people go back and look at their audit trail, uh- Mm-hmm … you know, just checking on even when humans are doing all the things. But looking at the loop, you know, the, the whatever harness you’re, you’re using and the loop system that [00:29:00] it’s, it’s checking to make sure that it actually did what you asked it to do correctly.
Mm-hmm. You know, to ex- to execute that, then it would keep going and going and passing between… That passing between your different parts of your agent, how you’ve designed it in order to make sure that’s correct before it does post to ERP or email or invoice or whatever it is Hmm. Okay. Yeah. Thank you, uh, Veronica.
That was a good one. Uh, anyone else wants to chime in? Dwight, you wanna say something?
So, uh, you know, for us, we, uh, you know, we-we’re, we’re, uh, very risk-averse, so we, uh, we are focusing on, uh, low-risk activities. Uh, and one of the reasons I like the, um, uh, the, um, benchmark you create is, uh, the fact that, uh, the, uh, the, the agent fails if it completes the transaction- Mm-hmm. -is super attractive to, [00:30:00] uh, a finance team who are generally risk-averse, so…
Yeah. Thank you. Uh, actually money movement, uh, is, uh, the riskiest part, uh, ’cause once the money goes to the different company, it’s hard to get back if it’s wrong.
Uh, anyone else wants to chime in before we go to the next slide? We’re still, trying to determine where we’re gonna let an agent run. So what is, uh,
your general thought process, Doug? Uh, I know you’ve been thinking very hard in, uh, last several months. Yeah. I think the sliced data, uh, where it’s controlled seems to be the, the route that may bring some, uh, you know… I mean, it’s all gonna bring value, don’t get me wrong. Mm-hmm. But I’m with Dwight on a very conservative move forward.
So- Mm. -we’re trying to figure out the best way that we can pull that data together within our systems, and we’re feeling pretty confident that [00:31:00] we’ll be able to pull it. Uh, our concerns- Mm. -will be more in the relation of making sure the data’s accurate, making sure the information’s being pulled, and that’s, you know, less is more, it seems.
Yeah, yeah. The, uh, MCP, right? The, the query tool, which only lets AI get a small portion of your whole data at a time. Yeah. Yeah. You’re, you’re really forward-thinking, uh, uh, and Doug, thank you, and Dwight and others here who’ve chimed in. Uh, so I’m gonna go to the next slide here.
So, uh
So if you want to actually deploy one of these and be able to defend it in an audit, here’s the four-step version. So number one, start bounded. One workflow, clear inputs, a clear definition of done. Resist the urge to boil the ocean. Uh, number two, put [00:32:00] the human on the exceptions. Let the agent grind the routine eighty percent.
Your people approve the unusual stuff and anything on a deadline. Number three, keep the audit trail. Every action logged, segregation of duty is fully intact. If you can’t reconstruct what happened, you’re not done. Uh, number four, prove it on your data before you scale. Measure accuracy, cost, and cycle time on your real documents.
The da-demo doesn’t count. Your books do
Uh, quick word on wh-why I get to say all this. Uh, we are not a general AI company that, uh, wandered into construction. We are certified on Trimble App Xchange with real-time two-way integration. We have integrations with major construction ERPs. Uh, we are built specifically for construction documents, uh, like AP invoices, G zero two, G zero three, pay apps, uh, delivery [00:33:00] tickets, credit card receipts, vendor statements, PO matching.
Uh, and we are live with specialty subs and GCs today, and your data stays yours and encrypted. So, and many of you I know from the podcast, Finance at the Jobsite, which is just me having these same conversations with the CFOs and other leaders in this room
So if you forget everything else, keep these three things. Uh, this is a today thing, not a someday thing. On bounded reviewable work, AP, migration, drafting, agents are reliable right now. Number two, keep humans on the clock. They break on time-sensitive, messy unattended work, and reliability collapses on repetition.
Remember, agent reliability does go down if you run the same tasks over and over. In fact, two-third stops to a third when you’re asked for five [00:34:00] in a row. And number three, open beats locked in. Demand evals on your own data and open pro-protocols. Own the control layer, and you keep your freedom to switch
That’s the honest version. Uh, thanks for spending, uh, last forty-five minutes. Now I’d rather hear from you than keep talking. We still got fifteen minutes. Questions, pushback, war stories, especially the war stories. Where has this burned you? Where are you curious? My email and the podcast are on the screen, and if you want to see the AP Work flow live, that link will take you there.
Who wants to go first?
Anyone wants to s- share a war story with us?
Come on. Dwight, you wanna share a war story?
Yeah, so I, we’ve, we’ve been doing a fair amount of experimentation. Uh, and like I said, we’ve been starting with really low, uh, risk tasks. Like a… The lowest risk task was, answering questions [00:35:00] from internal users. Mm-hmm. Uh, and for that, you know, we kind of ran that for a month or month or two and had really good, uh, response and freed up, , a lot of time where it’s like, “Oh, I don’t have to go get, actually get all the details.”
And so for us, it’s, it’s really not a, you know, a big reveal. We’re like taking small steps into, um, deploying AI agents. Um, and like I said, I, I, I’m not sure I will be alive when we actually allow a, a, uh, AI agent to con- complete a, uh, transaction. That just, that may be a bridge too far, and if, if so, that’s fine.
You know, if my team is, is basically focused on, going in, approving and executing, so much the better, right? That, that, that works for me. Um, but there is a level of trust that has to be earned. Mm-hmm. Like, ’cause like going in and saying, “Yeah, it’s gonna be fine. Just trust it.” Eh. Yeah. That really does not resonate.
People have to see it, and there has to be some kind of track [00:36:00] record. Um, because, you know, if you’re like, “Oh, well, we tried it five times and it worked fine,” it’s like, that’s not enough. Um, and the, the larger the transaction, the more risk there is, the more reticent people are gonna be. ‘Cause, you know, if something screws up, you know who’s gonna be cleaning it up, right?
It’s not the AI. Right. The AI’s not gonna clean up the mess, right? It’s gonna be the actual people in the, on, on the team who have to go clean it up. Um, so we’ve been very, very careful about not rushing it. Uh- Mm-hmm … and I’d rather get there and I’d rather take more time and get more comfort, uh, than get there quickly.
Um, so, uh, Rishi, one of the things I like about your style is that, you know, you’re not sh- you know, like, “Come on, let’s go do everything right now.” It’s like, no. Mm-hmm. Let’s take some low risk, uh, low, low risk activities and get those going and get some comfort, and then we can go in and see what, what do we do next.
But I really think that AI, um, agent adoption is gonna be, uh, it’s gonna take place over many months, [00:37:00] not many weeks. And I’ll, I’ll stop rambling there, Rishi. That was so well said. Uh, but, uh, again, construction finance generally, uh, we are a very conservative industry, and we have to be. This is, this is not some simple stuff we’re talking about.
This could be grave consequences for agent making mistakes here. Uh
Matt, uh, you wanna say something? You got my– I, I appreciate everything you’ve been… You know, you presented today. It’s got my wheels spinning. I have, you know, I have a few i-ideas that are, are, are being generated just from, from the discussion. It’s been helpful. But, you know, you just, um, i, I appreciated you pointing out the, uh, you know, the, the items that are, are repeatable.
Um, where there’s a kind of a set of users. You know, just simple things where we get an email from a, a vendor or a s- a customer that says, “Hey, [00:38:00] we need your, um… We need an updated,” whatever it be. Uh, a, like a, uh, form 300A from you. You know, we just- Mm-hmm. You know, we just go through this process. We go, we…
“Okay, thanks. We’ll attach that to an email, um, and send it on.” You know, it’s just– It’s, it’s a– To me, it’s a pretty low-risk thing, but that’s something that, you know, we could just have an agent pick up and, and look for those requests and, and, uh, you know, send that, send that over. We don’t mind our customers seeing that, so it’s really low risk.
And I think, I think just– that’s just one of many examples of things that, that are out there that we just gotta kinda concentrate and see what are all these things that, you know, we are, you know… Especially where we get a, a volume of, of these kinds of requests that we can be able to automate through agents.
[00:39:00] So anyway, appreciate, uh, uh, getting my juices flowing today. I’m glad. Uh, and, uh, one thing I really like about you is, like, you could think of things that I’m, like, not even aware of. So thank you so much, man. You’re welcome. Uh, Kevin. K-Kevin, you wanna say something? Kevin Jacobs?
Are you Kevin? Yeah. No, I, I tell you, um- Mm-hmm … Rishi, you– there, there are so many ideas that came out of this. Mm-hmm. But as you know, no, uh, I-I tell everyone on the call that, you know, I’ve recently started, uh, a new position with a new company. Mm-hmm. And we are still– they’re still in… I mean, they’re, they’re five years old now, but they’re still in startup mode as far as moving from all of their activity being captured in multiple Excel files [00:40:00] and, uh, and trying to get it– trying to get them to think in the mode of changing their process to using the ERP to its full potential, and then bringing in, uh, someone like you with your ideas to layer AI on top of it, and then make the processes even, even easier and, less error-prone once everything is worked out.
So, um, I’ve got a lot of ideas and, and I wanna talk to you about them, but right now I’m trying to get these contractors to think in, uh- Mm-hmm … ’cause you, you know, as well as everybody on here, that construction is about twenty years behind on technology implementation. Yeah. And, uh, I believe these guys are, are a little further behind than that.
So, uh, so I’ve got a lot of ideas, but I’m just– I’m drinking from a fire hose right now and trying to get to a point where I want– where you can come in, and it makes sense to you to, to initiate this AI, [00:41:00] um, these AI ideas. Yeah. Thank you, Kevin. Uh, yeah, I mean, co-construction does, does tend to be behind and, like forward, there’s a lot of forward-looking leaders here, but this industry is definitely, uh, I feel like it can need some, uh, some changes.
Yeah. Uh, I see Jayna. You wanna say? Jayna Bertholf, you wanna say something?
Uh, hi, Rishi. Yeah. Hi. Um, I don’t really have much to say. I’m mostly just listening to learn, Rishi. But yeah. We have just recently, um, decided to go with Beiing Human, and we’re excited to see what we can… gain from it, from the use of it. And, uh, thank you, Jayna. Mm. Anyone else wants to say something?
Alison, maybe. Alison Sawyer. Can I ask Jayna why she decided to go with Beiing [00:42:00] Human?
Of course, yeah I didn’t mean to put her on the spot again. I’m sorry. Yeah, no. Um, we’ve looked at a lot of different AP automation, um, solutions and, you know, actually, I think, uh, Rishi was one of the very first people we met with and, um, then we started looking at a lot of different things. And either, you know, some of them don’t integrate well with Foundation, which is what we use, and others of them were, um, very expensive or new, and we felt like there were going to be hang-ups and problems.
And, you know, I think the thing that finally threw it over the edge where we went with Beiing Human is, um, on the CFMA, um, cafe, you’ll see lots of [00:43:00] positive comments about Beiing Human, and I just don’t really see those about other companies out there. And so we just thought we’d take a chance with Rishi.
Awesome. Thank you. Thank you – Jayna, uh- Well, Rishi, if I could, if I could just add to that. Yeah. Um, I’ve been a, I have been an AP automation customer for over a year now, and it has been a real game changer for us. So, um, you know, so anyway, I just wanna throw that out there for anyone who’s, who’s thinking about that.
It’s, it’s really been a, a, a real, um, help for… Saved us, it saved us a lot of time. It’s improved accuracy. Um, and, uh, the getting, being able to get approvals done quickly, um, has been very, very, um… Well, Beiing Human has really helped us improve that a lot
Thank you, Matt. [00:44:00] Uh
We got another four minutes. Uh, anyone else wants to add something to this? Uh Yeah, I’ll throw something out there. We’re looking at trying to make sure we have our checksums compared to our data so that our process for financials or job costing or whatever that is, is gonna stay the same while we’re trying to implement some of these AI things.
And always trying, having that checksum built into the process so that it’s always referring back to that fixed non-hallucinogenic data. Um, and you know, that’s gonna be a long-term process, but I feel like that that’s gonna be key for us, uh, in at least in our business and our audit world. Uh, and we’re just starting with Rishi, although Rishi and I have known each other for a number of years.
Uh, we’re just starting the AP automation with him, and we’re looking really forward to it
Thank you, Doug. Yeah, uh, you know, the MCP and the sliced, uh, [00:45:00] amount giving to AI, that’s so important ’cause, the attention , budget of the AI falls, uh, rapidly if you’re not careful. Thank you so much, everyone. I will let, let you guys, uh, have your two minutes, and, uh, we’ll touch base here soon with, uh, a lot of you. Thank you, Rishi. I enjoyed your talk today. Thank you.
Yeah. Good, good discussion today. Thank you. Thank you. Thank you, Rishi. Thank you. Bye, everybody. Bye. Thank you, Kevin.
Thanks for listening to Finance at the Jobsite. If you found today’s conversation valuable, share it with a teammate and subscribe so you don’t miss the next episode. You can listen on Apple Podcasts or Spotify or watch on YouTube. Just search Finance at the Jobsite. Until next time, here’s to building smarter, faster, and more profitable projects