AI in the Finance Workflow

A practitioner's guide to what actually works, what breaks, and what should never be delegated — from someone who spent fifteen years doing the manual version.

The mechanical layer — automate with discipline
The judgment layer — human, permanently

Most of what is written about AI in finance is written by people who have never closed a book. The demos are impressive, the promises are large, and the workflows are imaginary.

This guide takes the opposite approach. I spent more than fifteen years inside finance organizations — running FP&A cycles, owning trial balances across entities, reconciling ERP data against consolidation layers, drafting the variance commentary that lands on a CFO's desk. I did, by hand and in Excel, almost everything AI now proposes to automate.

Knowing the manual reality of a workflow tells you precisely where a model adds value, where it merely adds noise — and where letting it act unsupervised would be professionally negligent. One color code runs through every graphic that follows: teal for what the model should do, amber for what a human must own. If you only skim the pictures, you'll still leave with the argument.

1

Variance Analysis & Commentary Drafting

Every finance professional knows the rhythm. Close ends, the numbers land, and someone has to explain them — actuals versus budget, versus forecast, versus prior year, by entity and product line. The analysis is half the work; writing it in language a business leader will read is the other half. In most teams I've led, commentary consumes two to four days of every cycle, much of it mechanical.

Where the cycle time goes — and what compression looks like
Commentary draft — manual 2–4 days Commentary draft — AI-assisted hours — draft (teal) + human verify (amber) Full reporting cycle — before 2 weeks Full reporting cycle — rebuilt 3 days · –70%
The lower pair is not hypothetical: it is the reporting-cycle compression I delivered in practice — from two weeks to three days, standardized within three months and sustained thereafter. Process redesign did that before AI existed in the workflow; AI now compresses the drafting layer on top of it.

The pattern that works

1 · Structure the input

Export a clean, well-labeled variance table. The model doesn't need your P&L — it needs a precise extract.

2 · Prompt for a draft

The model identifies the largest deltas and drafts neutral commentary — under strict constraints.

3 · Human adds judgment

The analyst confirms flagged drivers, kills misreads, and adds what the model cannot know: what management will do about it.

Prompt pattern — field-tested "You are drafting month-end variance commentary for a manufacturing business. Below is the variance table for [entity], actuals vs. forecast. Identify the five largest variances by absolute value. For each, draft two sentences of neutral commentary stating size and direction. Where the driver is not evident from the data, write [DRIVER — CONFIRM] instead of guessing. Do not speculate on causes. Use the terminology in the table; do not rename accounts."

The two constraints in bold are the entire game: don't guess drivers, and flag what needs confirmation. An unconstrained model invents plausible causes — "driven by higher freight costs" — that nobody verified. Plausible and wrong is the most dangerous combination in financial reporting.

2

From Commentary to Corrective Action

This is where FP&A either earns its seat or stays a reporting function. Commentary that only narrates the numbers is a rear-view mirror. The value is in what comes next: if this business unit is dropping, what compensates it — and how much do we need to recover, per month, to close the gap?

The mitigation logic follows a structure every experienced FP&A professional carries in their head: quantify the gap, identify the offset candidates, compute the required run-rate, test feasibility. The first three are mechanical — which means a model can draft them. The fourth is judgment, and it stays human.

The recovery bridge — turning a variance into a target
100 Margin plan –12 Volume — BU A –8 Freight & input costs 80 Current run-rate +20 ÷ 6 months ≈ +3.3 / mo Required recovery 100 Back to plan erosion — model quantifies recovery run-rate — model computes feasibility & decision — human
The bridge reframes a backward-looking variance as a forward-looking target: not "we lost 20 points of margin to volume and freight" but "we need +3.3 per month for six months, and here are the candidate sources." That single reframe is the difference between reporting and business partnering.

What the model drafts — and what it never decides

Quantify the gap

Size the erosion by driver — volume, mix, freight, input costs — and compute the run-rate needed to recover it over the remaining months. Pure arithmetic on governed data.

Draft offset candidates

Scan the portfolio for compensation options: products or units growing above plan, margin headroom, pricing levers, cost lines trending favorably. Proposals, explicitly labeled as such.

Judge & decide

Whether the growing product line can actually absorb more volume, whether a price move holds commercially, which action management commits to — feasibility and decision stay human.

Prompt pattern — mitigation drafting "Below are the YTD variance table and the current forecast by product line for [segment]. The margin gap to plan is [X]. 1) Compute the monthly recovery run-rate required to close the gap by year-end. 2) Identify up to five candidate offsets from lines performing above plan, stating for each the incremental contribution needed and the implied volume or price change. Label every candidate as a proposal requiring commercial validation. Do not assess feasibility — flag it as [MANAGEMENT — VALIDATE]."

Notice what the constraint does: the model produces the options and the math — the part that used to consume an analyst's afternoon — and explicitly refuses the feasibility call. Whether BU B's growth can really be accelerated, whether the freight increase is structural or spot, whether a price action survives contact with customers: that judgment is why FP&A has a seat at the table. The model prepares the decision. It does not make it.

FP&A adds value when it walks into the review with the gap quantified, the recovery math done, and three candidate actions on the table — not when it tells the story of the numbers one more time.
3

Scenario Modeling with LLMs

Scenario work in most organizations is spreadsheet archaeology. Someone builds a driver model, someone inherits it, and by year three nobody fully trusts the links. When leadership asks what happens to margin if input costs rise eight percent and volume softens, the true answer arrives days later.

I use LLMs in three distinct roles here — and keeping them separate is the discipline:

Structuring partner

Before building: "Here are the drivers I plan to model. What am I missing? What second-order effects link these?" The model reliably surfaces the interactions a tired analyst skips.

Scenario narrator

After the numbers exist in a governed model: hand it the output of three scenarios and ask for the comparative story — which driver dominates, where sensitivity concentrates.

Sanity checker

"What would a skeptical CFO challenge in this forecast?" Five minutes. It has caught hockey-stick second halves and stale FX assumptions in real reviews.

What I deliberately do not do: let the model perform the arithmetic of record. LLMs approximate computation, and they do it confidently. The numbers live in a governed system — the model works around them, never as their source.
4

Automating Reporting Narratives

For years I built monthly business review packages — the layer where consolidated data becomes a story leadership can act on in a ten-minute read. High-value work with a highly repetitive core: every month, the same structure, new data, a fresh narrative. The narrative is where the hours go — which makes it the single most automatable text in the finance calendar.

The narrative assembly line — same inputs, every cycle
Fixed narrative template Consolidated summary table Confirmed driver notes Last month's narrative Model drafts continuity + [CONFIRM] flags Human edits confirms flags · owns voice Published MBR narrative
Feeding last month's narrative back in is the instruction that matters most: executive readers notice when the storyline resets every cycle. Continuity is what even strong teams lose under close-week pressure — and what a template-driven model preserves effortlessly.

The honest result is qualitative as much as quantitative: when producing the narrative stops consuming the team's energy, the content improves — because people finally have time to think about what the numbers mean before writing about them. Automation here doesn't replace the analysis. It funds it.

5

The Use-Case Map: Where AI Lands First

Beyond FP&A, four operational areas dominate my map of where AI creates value fastest — ranked by conviction earned the hard way: I have personally been the manual version of each.

Impact vs. readiness — a practitioner's placement
Readiness to automate today → Business impact → start here 1 Reconciliation intelligence detect · classify · explain — human approves the fix 2 Mapping validation where EPM migrations quietly go wrong 3 Forecast intelligence extend the reviewer, not the spreadsheet 4 Close orchestration coordination, not judgment — the safest first step
Why reconciliation ranks first: for two years, part of my role was resolving inconsistencies between source ERPs and the consolidation layer. I was, in effect, a human reconciliation agent — detect the break, trace it, explain it, propose the fix. That pattern-matching is learnable structure, and models handle it well. I know what "good" looks like because I was the benchmark.

Mapping validation deserves a special word because it is unglamorous and extremely valuable: anyone who has migrated a consolidation platform knows account mapping is where migrations quietly go wrong. A single bad mapping surfaces weeks later as an unexplainable variance. AI-assisted match suggestion and anomaly flagging — "this account maps differently from 94% of similarly named accounts" — would have materially reduced risk in every migration I've been part of.

6

The Controls Lens: What Never Gets Delegated

This is the section most AI-in-finance content omits — and the one that determines whether any of the above survives contact with an audit. I come from a controls background: external audit, then designing SOX controls inside operating companies. That makes me conservative in specific places, and the conservatism is a feature.

The delegation boundary — the whole argument in one table
Automate freely
  • Detecting breaks and anomalies
  • Classifying and explaining variances
  • Drafting commentary and narratives
  • Suggesting mappings and matches
  • Flagging what needs human review
A human owns the click
  • Posting, adjusting, correcting entries
  • Confirming drivers and causes
  • Signing commentary as management's view
  • Approving any change to the books
  • Judging the model's confidence
The test is simple: if the action would change what the books say, a person with accountability approves it. Detection automated; correction approved.
Auditability is non-negotiable

AI-generated text about numbers is drafting. AI-generated numbers outside a governed system are a finding waiting to happen.

Confidence is not evidence

Models present wrong answers with the same fluency as right ones. "Flag, don't guess" belongs in every prompt; verification belongs to the reviewer.

Segregation of duties survives

If one person prompts, approves, and posts, three controls just collapsed into one seat. Automated workflows need the same duty split the manual ones had.

None of this is resistance to AI. The organizations that will scale AI in finance furthest are the ones whose controls let them trust it. Guardrails aren't the brake — they're the reason you're allowed to drive fast.

The pattern underneath all of it

AI compresses the mechanical layer of finance work and returns the time to the judgment layer. Drafting, detecting, classifying, narrating — automate with discipline. Interpreting, deciding, confirming, signing — keep human, permanently.

"Finance teams don't need another tool demo. They need the workflow rebuilt by someone who knows where the hours go, where the errors hide, and where the auditor will look."

J. Sebastian Bassi has spent 15+ years in international finance across KPMG, IBM and global industry — leading FP&A, consolidation and enterprise performance management transformation, including an HFM-to-OneStream migration across ten manufacturing entities. He holds IBM's Generative AI certification and writes about finance transformation and AI-enabled finance at bassiconsulting.com.
← Back to Research & Insights