ChatGPT for Tax Workflows: How Tax Teams Are Doing It in 2026
Suhail Ameen
September 21, 2026
13 Min Read

See how CFOs and controllers deploy AI that survives an audit. Watch Now→
See AI that passes audit + a demo of pre-built agents for finance. Watch Now →
How To Build an AI-Ready Finance Function. Read the Guide →
From AI chat to governed, audit-ready workflows. Find out what it takes to run AI inside close, reconciliation, and tax processes.
Watch Now
The 2026 Tax Leader Decision Map: Make decisive, defensible decisions in an evolving tax environment
Read the E-Book
AI can draft, summarize, and analyze. Savant CEO Chitrang Shah explains what it takes to close the AI trust gap in finance.
Read Now
A tax manager at a company with thirty-odd legal entities is three days into the Q3 provision. She asks ChatGPT how a state treats a particular interest addback, gets a clear answer in ninety seconds, and drops it into her memo. Then she spends the rest of the week pulling trial balances out of two ERP instances and waiting on the fixed asset register. She also rebuilds the apportionment workbook, because two entities changed their account structure in August. The model saved her an hour. The rest of the week went the way it always does.
That gap is the story of AI in tax right now. The Thomson Reuters Institute’s 2026 AI in Professional Services Report found that 86% of tax professionals who use generative AI reach for it at least once a week, and 36% use it multiple times a day. Looking ahead, 88% expect AI to be central to their workflows by 2030.
One number in that report explains the gap better than the rest. Roughly 65% of tax professionals still run their AI work through public tools like ChatGPT, while only 34% have moved to tools built specifically for tax work.
ChatGPT is good at the thinking around tax work. It is not built to produce the numbers.
That distinction sounds pedantic until you apply it. ChatGPT explains an interest limitation clearly, drafts a defensible first response to a state notice, and summarizes 200 pages of new guidance in a minute. It does none of those things as the system of record. It carries no professional liability, holds no license, and produces no audit trail.
The professionals have already sorted this out in practice. The Thomson Reuters data shows four leading applications: tax research at 69%, document summarization at 57%, return preparation support at 55%, and advisory work at 55%. Every one of those describes a professional using AI to move faster through a judgment they still own.
No adoption survey lists producing the provision as a use case, and the reason is not capability, but accountability. The tax director signs off, the number lands in the financial statements, and external audit asks how it was built. A model working alone answers none of those questions.
The strongest use cases share a pattern. A professional supplies the context, the model accelerates the drafting or the reading, and the professional reviews the output before it reaches a workpaper.
Tax research leads every adoption survey for a reason. Reading Treasury regulations, comparing how three states treat the same transaction, and building an initial research framework used to consume hours. ChatGPT compresses the first pass. It does not replace verification against primary sources, and treating its answer as authority is how departments get into trouble.
Tax teams read 10-Ks, partnership agreements, intercompany agreements, and a steady stream of guidance updates and state notices. Summarization is the single highest-leverage use of a general assistant, because the work is reading rather than deciding.
Technical memos, disclosure language, and plain-English explanations of a position for the business all draft faster with AI. The reviewer still edits for accuracy and tone, but the blank page disappears.
Provision procedures, review checklists, onboarding materials, and training content all come out of ChatGPT quickly. This use case carries almost no risk, because it involves no company financial data.
Before modeling an entity restructuring or a timing election, a tax professional can use ChatGPT to enumerate the alternatives worth analyzing. The model widens the option set, and the professional narrows it down with judgment.
The ceiling arrives faster than the adoption numbers suggest. Below are four specific ways that happens.
Return preparation requires current-year forms, current-year rules, jurisdiction-specific calculations, e-file transmission, diagnostics, and a preparer signature backed by professional liability. ChatGPT connects to none of the required infrastructure. It cannot pull a prior-year return, carry forward a basis schedule, run diagnostics against IRS business rules, or transmit anything.
It also cannot calculate reliably against live data. Large language models are probabilistic by design, so ChatGPT produces results that look plausible without being reproducible. Verifying that output line by line costs more time than running the calculation in purpose-built software. Compliance software owns this work, and for good reason.
A department covering 30 entities and 45 state filings needs to know which entities have closed and which apportionment factors are final. It also needs to know which returns are still waiting on a business unit that hasn’t yet sent its data.
ChatGPT holds no state across conversations in any way a tax department can rely on. It doesn’t know your entity structure, your filing calendar, your reviewers, or which version of the workbook is current. Every session starts empty.
The practical cost of this shows up as context switching. An analyst generates something useful in ChatGPT, pastes it into a workpaper, pulls data from three other systems by hand, and updates a status somewhere else. Each handoff burns part of the time the AI just saved.
This limitation surprises people, because reading a PDF looks like exactly what a modern model should be good at. But production extraction in tax demands more than just reading. It requires consistent mapping across hundreds of issuer-specific layouts, confidence scoring on every extracted value, exception routing when confidence drops, and delivery into the system that consumes the data. A K-1 alone can carry dozens of relevant boxes, with footnotes that change what those boxes mean, while an invoice set spanning four countries carries different tax fields in different positions.
Accuracy at volume matters more than accuracy on one document. A 95% extraction rate sounds strong until you process 40,000 fields in a quarter and inherit 2,000 errors with nothing flagging any of them. This is the same problem our evaluation of Claude Cowork in enterprise finance workflows ran into: capability alone isn’t the same as operationalization.
General-purpose AI tools generate no audit trail. They offer no role-based access control, no data residency guarantee, no retention policy you control, and no way to demonstrate to a regulator which company data reached which model and why. This is not a hypothetical concern, and results in significant legal exposure.
Tax work feeds financial statements, which means every number in it is a SOX control question and an audit inquiry waiting to happen. LLMs like ChatGPT create proof debt, where the work gets done without the documentation and traceability needed to defend it. What matters is lineage from source data to final output, not a log of the questions someone typed into an assistant. The same standard applies to AI agents inside a governed workflow, which essentially operate as a privileged user.
Ask ChatGPT or another general assistant to run this month’s bank reconciliation, and it has to ask you what you mean, because reconciliation is not one standard universal procedure. Your team matches on amount and date inside a three-day window. It tolerates a variance up to a set dollar figure, writes off differences below materiality, and routes anything above it to a named reviewer. The competing company down the road does all of that differently, and both of you are right.
That company-specific procedure is part of your control environment. It lives in a senior analyst’s head, in a procedures document nobody opens, and in the muscle memory of having closed thirty such reconciliations. It’s why a new hire takes two quarters to become useful. And it’s exactly what a general AI assistant cannot know, because it is not in the training data or the prompt.
The gap is memory and repeatability. The assistant has no record of how your company does the work, and no way to perform it the same way twice. Those two things are what turn a good answer into finished work.
Savant is the control layer that supplies both institutional knowledge and deterministic execution. Your team describes the work in the model it already uses, whether that is Claude, ChatGPT, Copilot, or a private LLM. Savant turns that conversation into a deterministic workflow and holds the institutional knowledge layer behind it: your match tolerances, your date windows, your materiality thresholds, your exception criteria. Savant then runs the workflow against live data from your ERP, the same way next month and the month after.
The division of labor is the part a Controller should care about. ChatGPT (or your LLM of choice) interprets the request. It never performs the work. Execution sits in Savant, which is what makes the result reproducible and the run reviewable. A probabilistic system helps define the process, and a deterministic one carries it out.
That’s the difference between an assistant that drafts a memo about reconciliation and a process that completes the reconciliation. With Savant, data lineage, human review checkpoints, and compliance reporting are built-in. Every run is audit-ready by default.
Under the surface, the data work has to actually happen, and this is the part most organizations haven’t automated. Tax teams pull from ERP systems, general ledgers, subledgers, payroll platforms, fixed asset registers, and a long tail of spreadsheets, each arriving in a different structure on a different cadence.
Savant’s Vision Agent reads the unstructured half of that, pulling values out of K-1s, invoices, contracts, customs documents, and bank statements. The Fuse Agent then matches records on meaning and context rather than on exact names or IDs. Inconsistent vendor names and reference formats stop breaking a reconciliation the way they do under a rules-based match.
Account structures get standardized in the same pass. When two entities keep different charts of accounts, Savant maps them to a common structure instead of leaving an analyst to rebuild the crosswalk each quarter.
Then it delivers. Savant’s connectors run in both directions: read and write. That means pulling transaction data from platforms like SAP, Oracle, and NetSuite, and posting the resulting calculations, adjustments, and entries back into the ERP, rather than handing someone a file to upload. Page-level anchoring ties every calculation back to its source document, approvals and role-based permissions sit inside the workflow, and evidence packs export by period.
Two boundaries are worth stating plainly:
The value scales with scale and complexity, because a manual reconciliation that holds together across three entities breaks across thirty. Rover rebuilt its monthly tax process this way and cut its close by 80%, with transaction matching running on a schedule and outputs audit-ready by default.
The teams furthest along the tax automation journey have stopped asking which single tool handles tax work and started assigning each system a job it can hold.
The ERP and subledgers hold the record. They are the source of truth for transactions, balances, and entity structure, and nothing downstream should quietly disagree with them.
Provision and compliance software owns the filing. The statutory forms, the transmission, and the sign-off live there, and no general assistant displaces them.
Tax engines determine transaction-level tax. Rate lookups and jurisdiction logic belong in systems built to keep up with rule changes.
General AI is where the work gets described. On its own, it handles research, drafting, and summarization. When connected to a governed control layer, it becomes the place your team asks for real work in plain language.
The control layer holds your process and runs it. It turns those requests into deterministic workflows, standardizes the inputs, and executes the procedure your team configured. Then, it posts results back into the systems above and records what happened at each step.
Enterprise agreements from the major LLM providers typically exclude customer data from model training and add SSO, role-based access, and audit logging. Consumer accounts do none of that. This single decision resolves a large part of the governance problem.
The governance structure already exists in an enterprise, so use it. A tool that has cleared security review carries far less risk than one a director adopted independently, and it converts shadow AI into sanctioned, governed AI with a name on it.
This is the step teams skip, and it’s the one that determines whether any of this works. Your match tolerances, materiality thresholds, and exception routing may live only in the head of the person who has always done the close. No assistant and no automation layer can execute what it cannot read. Documenting the procedure is worth doing even if you automate nothing.
Provision work feeds the financial statements, so it needs traceability from source data to final number with approvals recorded. That does not mean banning the conversational interface. It means making sure what the interface triggers is a governed workflow rather than an improvised, probabilistic one.
ChatGPT is a useful tool for tax professionals, and the adoption figures reflect real value rather than hype. It accelerates research, compresses document review, drafts memos, and standardizes internal procedures. Departments that avoid it entirely give up an advantage for no good reason.
However, if used on its own, it is not a tax system. It doesn’t produce the provision, extract source documents at production accuracy, coordinate a multi-entity filing calendar, or provide the governance a SOX environment requires. Those gaps are architectural, so no matter how smart a model gets, it will not close them.
It doesn’t need to become a tax system, though. Connecting it to one is the whole job. Give your LLM a layer that knows how the company works and executes tax processes the same way every time. This is where Savant works. Your team keeps the assistant it already likes, and asking for this month’s reconciliation produces this month’s reconciliation, done your way, with the evidence attached.
Looking for a concrete starting point? Our guide to tax workflows you can automate with AI in 5 days lists the candidates most teams work through first.
Savant is dedicated to helping businesses of all sizes make informed decisions. We adhere to strict editorial guidelines to ensure that our content meets and maintains our high standards.