A practitioner's guide to what actually works, what breaks, and what should never be delegated — from someone who spent fifteen years doing the manual version.
Most of what is written about AI in finance is written by people who have never closed a book. The demos are impressive, the promises are large, and the workflows are imaginary.
This guide takes the opposite approach. I spent more than fifteen years inside finance organizations — running FP&A cycles, owning trial balances across entities, reconciling ERP data against consolidation layers, drafting the variance commentary that lands on a CFO's desk. I did, by hand and in Excel, almost everything AI now proposes to automate.
Knowing the manual reality of a workflow tells you precisely where a model adds value, where it merely adds noise — and where letting it act unsupervised would be professionally negligent. One color code runs through every graphic that follows: teal for what the model should do, amber for what a human must own. If you only skim the pictures, you'll still leave with the argument.
Every finance professional knows the rhythm. Close ends, the numbers land, and someone has to explain them — actuals versus budget, versus forecast, versus prior year, by entity and product line. The analysis is half the work; writing it in language a business leader will read is the other half. In most teams I've led, commentary consumes two to four days of every cycle, much of it mechanical.
Export a clean, well-labeled variance table. The model doesn't need your P&L — it needs a precise extract.
The model identifies the largest deltas and drafts neutral commentary — under strict constraints.
The analyst confirms flagged drivers, kills misreads, and adds what the model cannot know: what management will do about it.
The two constraints in bold are the entire game: don't guess drivers, and flag what needs confirmation. An unconstrained model invents plausible causes — "driven by higher freight costs" — that nobody verified. Plausible and wrong is the most dangerous combination in financial reporting.
This is where FP&A either earns its seat or stays a reporting function. Commentary that only narrates the numbers is a rear-view mirror. The value is in what comes next: if this business unit is dropping, what compensates it — and how much do we need to recover, per month, to close the gap?
The mitigation logic follows a structure every experienced FP&A professional carries in their head: quantify the gap, identify the offset candidates, compute the required run-rate, test feasibility. The first three are mechanical — which means a model can draft them. The fourth is judgment, and it stays human.
Size the erosion by driver — volume, mix, freight, input costs — and compute the run-rate needed to recover it over the remaining months. Pure arithmetic on governed data.
Scan the portfolio for compensation options: products or units growing above plan, margin headroom, pricing levers, cost lines trending favorably. Proposals, explicitly labeled as such.
Whether the growing product line can actually absorb more volume, whether a price move holds commercially, which action management commits to — feasibility and decision stay human.
Notice what the constraint does: the model produces the options and the math — the part that used to consume an analyst's afternoon — and explicitly refuses the feasibility call. Whether BU B's growth can really be accelerated, whether the freight increase is structural or spot, whether a price action survives contact with customers: that judgment is why FP&A has a seat at the table. The model prepares the decision. It does not make it.
Scenario work in most organizations is spreadsheet archaeology. Someone builds a driver model, someone inherits it, and by year three nobody fully trusts the links. When leadership asks what happens to margin if input costs rise eight percent and volume softens, the true answer arrives days later.
I use LLMs in three distinct roles here — and keeping them separate is the discipline:
Before building: "Here are the drivers I plan to model. What am I missing? What second-order effects link these?" The model reliably surfaces the interactions a tired analyst skips.
After the numbers exist in a governed model: hand it the output of three scenarios and ask for the comparative story — which driver dominates, where sensitivity concentrates.
"What would a skeptical CFO challenge in this forecast?" Five minutes. It has caught hockey-stick second halves and stale FX assumptions in real reviews.
For years I built monthly business review packages — the layer where consolidated data becomes a story leadership can act on in a ten-minute read. High-value work with a highly repetitive core: every month, the same structure, new data, a fresh narrative. The narrative is where the hours go — which makes it the single most automatable text in the finance calendar.
The honest result is qualitative as much as quantitative: when producing the narrative stops consuming the team's energy, the content improves — because people finally have time to think about what the numbers mean before writing about them. Automation here doesn't replace the analysis. It funds it.
Beyond FP&A, four operational areas dominate my map of where AI creates value fastest — ranked by conviction earned the hard way: I have personally been the manual version of each.
Mapping validation deserves a special word because it is unglamorous and extremely valuable: anyone who has migrated a consolidation platform knows account mapping is where migrations quietly go wrong. A single bad mapping surfaces weeks later as an unexplainable variance. AI-assisted match suggestion and anomaly flagging — "this account maps differently from 94% of similarly named accounts" — would have materially reduced risk in every migration I've been part of.
This is the section most AI-in-finance content omits — and the one that determines whether any of the above survives contact with an audit. I come from a controls background: external audit, then designing SOX controls inside operating companies. That makes me conservative in specific places, and the conservatism is a feature.
AI-generated text about numbers is drafting. AI-generated numbers outside a governed system are a finding waiting to happen.
Models present wrong answers with the same fluency as right ones. "Flag, don't guess" belongs in every prompt; verification belongs to the reviewer.
If one person prompts, approves, and posts, three controls just collapsed into one seat. Automated workflows need the same duty split the manual ones had.
AI compresses the mechanical layer of finance work and returns the time to the judgment layer. Drafting, detecting, classifying, narrating — automate with discipline. Interpreting, deciding, confirming, signing — keep human, permanently.
"Finance teams don't need another tool demo. They need the workflow rebuilt by someone who knows where the hours go, where the errors hide, and where the auditor will look."
← Back to Research & Insights