What AI gets wrong in a month-end close if you do not constrain it.
These are not lab edge cases. They are the confident wrong answers we have watched a close SOP produce on live QuickBooks Online files, including our own book at AG Accounting, when the prompt was too polite to stop.
Constraining is the job. Cheerleading is the default.
Unconstrained Claude will close your month. It will also be sure about it. That combination is the problem. If you are a fractional CFO about to audit the books before you model, it is also how a 13-week cash view gets built on a plugged recon.
These are not lab edge cases. They are the confident wrong answers we have watched a close SOP produce on live QuickBooks Online files, including our own book at AG Accounting, when the prompt was too polite to stop.
Leave the model unconstrained and a month-end close in Claude will quietly pick cash or accrual from habit instead of the client file, plug a recon difference so the report ties, compare this month to a trailing trial balance that cannot see last month’s restatement, invent timestamps it cannot read, dump unidentified card payoffs into equity, cheerlead a messy file, reformat your Excel package until the template breaks, and invent treatment for a vendor it Googled. Human approval before writes does not catch any of that if the proposal looks tidy. Constraints in the firm section and the client file. Not a better vibe.
The architecture is firm instructions. This is the failure list we keep next to the close.
Language models are trained to be helpful. Helpful, in a close, looks like a finished package with no red text. A reviewer needs the opposite.
A person who has closed books for a decade stalls when a statement is missing, when equity moved for no reason, when cash-basis revenue showed up with an AR balance. A model does not stall unless you order it to. It completes the pattern. Completing the pattern is how you get a recon that “ties” because the difference landed in opening-balance equity, where nobody looks until the tax return. A fractional CFO who inherited that file and then modeled cash will look at the tax return too, later, and it will not be fun.
I built this inside a firm that still signs the work. We do not need the model to be inspiring. We need it to stop. Every item below is a place we had to write “stop” in plain language after watching it not stop.
None of this is an argument against Claude or ChatGPT on the books. It is an argument against a generic checklist and a smile. Gemini is not in this conversation: no MCP, so it is not running your close through this connector. QuickBooks Desktop and ProConnect Tax are also not in this conversation. Those gaps are real. Do not pretend a QBO close SOP applies there.
Cash versus accrual is a file fact, not a personality.
The model has a default. If you did not write the basis down, the default is whatever showed up most often in training and in the last ten chats. Usually accrual language. Sometimes cash, if the prompt said “small business.”
On a cash-basis client, that default books AR, AP, and prepaid that do not belong. On an accrual client, it skips the prepaid and the deferred revenue because the bank rec looked clean. Both files produce a P&L. Both files are wrong in a way that survives a skim. A loud error is easy. A quiet one gets signed.
Write it down:
- Accounting basis on the client file, not in the chat. Cash, accrual, or modified, in those words.
- If basis is missing, stop. Do not infer it from an AR account on the chart. Charts outlive basis changes.
- Revenue recognition rules that actually apply to that client (deposits, retainers, prepayments, job deposits). “Use professional judgment” is how you get ASC 606 language on a cash-basis landscaper.
We watched a close SOP do this: the procedure talked like accrual, the QBO company was cash, and the model invented accruals the firm would never have booked so the two would agree. The package looked like a real close. The basis was fiction.
Plugging a recon difference hides the missing entry.
Ask an unconstrained model to “reconcile the bank to the statement” and it will get the difference to zero. That is the assignment it heard. Zero is available. The plug is usually an owner draw or contribution, opening-balance equity, a suspense account it created, or a round JE to cash dated the last day of the month.
The rec report looks done. None of those is a rec. A rec is a line-by-line match that names the missing deposit, the duplicate expense, or the feed that skipped four charges. If you cannot name the difference, you are not finished. The proposal should say that, in those words, and wait.
Write the rule so it cannot be misunderstood:
- Never post a reconciling JE whose only purpose is to force book-to-bank to zero.
- Never create a suspense account in the close.
- If the statement is missing, say “GL only, no rec” and do not produce a rec report.
- If the difference is under materiality, still name it. Materiality is for whether you chase it this month, not for whether you invent a story.
The QBO API will not mark the rec done for you anyway. The connector can produce a rec report. It cannot tick the box in QuickBooks. That limit is a gift. A person still has to look. Do not waste the look on a plug. Bank-feed controls, if you use them, are onthe bank feed page. They are not a recon.
A trailing trial balance cannot see last month’s restatement.
Quietest failure. Wrecks comparative packages. Wrecks a 13-week cash forecast that treated last month as gospel.
The model pulls a trial balance as of today, or a trailing twelve, and compares “this month” to “last month” from that live pull. Someone changed last month’s revenue after you issued the statements (late bill, reversed accrual, class recode, a journal a partner dumped in). The live trailing view has already absorbed it. The model reports that nothing moved. The client’s PDF from last month disagrees. You find out when they forward both PDFs. That email is never fun.
The fix is operational, not clever:
- Save the issued balance sheet and P&L from last month. Attach them or save them as reports. Live pull is not an archive.
- Compare this close to that issued pack.
- If last month’s pack is missing, stop. Reconstructing “what we probably sent” is how restatements hide.
We documented this as a confident wrong answer on a real SOP. The trailing TB path cannot detect a prior-month number that moved. Once you have seen it, you cannot unsee it. Storage rules for saved reports are on client reporting.
Invented timestamps are not an audit log.
Claude does not have a clock your workpapers can trust. Ask when something posted, or ask it to stamp a close report with “reviewed at 4:12pm,” and it will produce a time. That time is not from QBO. It is not from the connector unless you pulled it from the activity log.
Every MCP tool call lands in an activity log. You do not have to take the model’s word for when a tool ran. Use the log. If the SOP asks for a timestamp the model cannot read, the SOP should say “pull from the activity log or omit.” Invented times make a workpaper you cannot reproduce.
Same family: invented invoice numbers, invented check numbers, invented “per conversation with the client on Tuesday.” If it did not come from QBO, from a file you attached, or from a log you can open, it does not go on the close report. Close week already has enough fiction.
Unidentified card payoffs do not go to equity.
A credit-card payoff the model cannot match (wrong date, two cards, a personal card mixed in, a balance transfer) is a held item. Unconstrained, the model still “finishes” the close. The leftover lands in owner equity or owner draw, a clearing account that never clears, or the original expense account, which double-counts the spend.
Equity is the favorite because it makes the P&L look right. That is the tell. If the P&L got cleaner and the balance sheet grew a mystery, you did not close. You hid a card. An outsourced CFO who then reports “cash is fine” off that P&L has just signed the hide.
The bank-feed product already treats unidentified card payoffs as a structural hold. The close SOP should say the same thing in your words: do not code it, do not equity it, flag it, wait. Bank feed controls list the holds that never lift. Copy the spirit into the close, even if you are not using the feed that month.
Cheerleading is a buying objection, not a tone issue.
Firms add a guideline to skip flattery and then look slightly embarrassed about it. They should not. A close report that says “overall the books look great this month” after three months of losses is a liability. The partner skims the adjective and misses the flag. A board pack with the same sentence is worse. I have done that. You have done that. Adjectives are camouflage.
Write it down:
- No praise. No “great job.” No “the books are in good shape” unless the exception list is empty and you named the tests.
- Lead with exceptions. If there are none, say “no flags under the written tests,” then the package.
- Do not use confidence language (“high confidence this is correct”). Confidence is not a control. Approval is a control.
Cheerleading is also how templates die. A model that wants to be helpful will add a cover paragraph, restyle a header, and “clean up” your column order. Next failure.
Reformatting will break the Excel package you already sold the client.
Controller teams who already have a month-end workbook (P&L, BS, analysis tabs for vendor-by-GL, marketing, recruiting, commissions) will ask Claude to fill it. Fractional CFOs do this with the 13-week tab and the WIP-by-project tab. Unconstrained, Claude rebuilds the workbook. Different column widths. Different tab names. A merged header. Your vlookup dies. Next month, a third layout. The close that used to take a long day now takes a long day plus template repair. Nobody budgets for template repair.
Rules that have worked:
- Fill values. Do not change structure. Say that twice.
- If you cannot write into the existing tabs without breaking them, output a flat CSV or a values-only block and let a person paste.
- Never invent a vendor-by-GL report as a new design if the client already has a tab for it. Match the tab.
- The connector will not ingest a Word close template with row shading. Do not ask it to “use Close_Report_TEMPLATE.docx.” Produce the numbers, then drop them into the template the firm already owns.
Branded statements from the live ledger, in the firm’s voice, are a different path:client reporting. That path still does not get to “fix” your historic Excel. If the client wants a stoplight package instead of a workbook, that is a product conversation, not a reformat.
Invented vendor treatment is not research.
Should the prompt tell Claude to search the web when a vendor is new, so it knows whether this is software, a contractor, or a merchant processor? People ask this in good faith.
Search can tell you what Stripe is. It cannot tell you how this client books Stripe, whether this contractor is 1099-eligible on this file, or what the capitalization policy is. Unconstrained search plus “treat it appropriately” is how you get a software subscription booked to office supplies (or the reverse), a contractor booked to payroll expense, a merchant fee netted against revenue on a client who needs gross, or a 1099 vendor with no TIN and no flag, which you will meet again in January.
New vendor: follow the firm’s new-vendor SOP, or stop. Set the 1099 flag and the TIN when the vendor is created, not when you “detect” 1099s in January. Theprompt library has the vendor-scan asks we actually send. They flag. They do not silently recode a policy.
Audit the books before you model, or the 13-week inherits the plug.
This is the fractional CFO version of every failure above. You did not keep these books. You are about to stake a forecast, a board pack, or a WIP schedule on them. Unconstrained close is how the mess becomes your mess.
The order is not optional:
- 1. Audit: unexplained equity, basis mismatches, rec plugs, trailing numbers that do not match last month’s issued pack, unidentified card payoffs.
- 2. Close, with the stops in this post written into the firm section.
- 3. Only then: 13-week cash, budget-to-actual, WIP by project, the model.
A 13-week cash forecast built on a plugged recon is worse than a messy P&L. The P&L at least looks like a P&L. The forecast looks like a decision. Human approval before writes still matters here. You are often not the bookkeeper. You still should not let a model mutate a file you just agreed to trust.
What a constrained close looks like instead.
Same model. Failures above written in as stops. The artifact gets shorter, uglier, and more useful:
- Basis, materiality, capitalization, last month’s pack, skill version: confirmed or halted.
- Uncategorized and rec: named differences, no plugs.
- AJEs proposed, not posted.
- Flags first: negative BS, three months of losses, debt spike, owner draw, unidentified card payoff.
- Package in the existing template or the firm’s branded statements, structure untouched.
- Unknown vendors listed, not invented.
- No timestamps the log cannot support.
- A person approves writes.
- Forecasts, if any, start after the audit, not instead of it.
Slower than a hallucinated clean close. Faster than a real close done by hand. The only version I will put my firm’s name on.
For the monthly runbook, use the companion how-to in this series, or the sequence in theprompt library. For what the connector will refuse even if you ask, read the cannot-do post in this series and features.
Review-discipline questions, answered.
Book a walkthrough and break a real SOP.
Bring the close prompt you use today. We will run it against a live file and mark where it would have produced a confident wrong answer.