The obvious answer is that they handle one file well and 400 badly, and that the fix is something that runs on a schedule. That answer is right up to a point, but it skips the part where implementation teams first run into trouble.
The harder problem shows up in a single file, with a person sitting there and paying full attention.
You give a client’s export to Copilot, Claude, or ChatGPT and ask what’s wrong with it. Within seconds, you get a good answer: 47 problems, described clearly, grouped sensibly, each with an explanation of why it matters. The analysis is often better than most people would manage by hand.
Then you have to fix 47 things, and all you have is a paragraph about them.
What comes back reads something like this:
“I found 47 issues. 12 rows have invalid dates in column D (Start Date). 8 rows are missing Employee ID. 3 rows contain duplicate identifiers. 24 rows have a department name that isn’t in your reference list.”
As analysis, that’s close to perfect. As a starting point for two hours of remediation work, it’s close to useless, because summarizing is the wrong operation for the job.
The mechanics wear people down:
Claude for Excel improves on this with clickable citations back to the cells it references, which helps. It’s still a citation rather than a worklist: no state, no progress, and no way to say that these 11 are handled and that one is a question for the client.
Fix five things and ask again, and you get a new analysis of a changed file: a fresh summary, possibly organized differently, with row numbers that may have moved. The assistant does exactly what it should. It just has no concept of what you’re doing – working through a list.
Getting a summary rather than a list carries a subtler cost.
“12 rows have invalid dates” is a judgment the assistant has already made on your behalf. Whether those 12 are broken depends on something it doesn’t know: what your process will accept.
If all 12 are 31.03.2026, that isn’t bad data – it’s a European date format, and the right fix is to accept the format rather than edit 12 rows. If they’re 12 different kinds of wrong, that’s a conversation with the client.
You can’t tell which situation you’re in from the summary. You have to look at the values, so the summary has cost you a step rather than saving you one.
And once you’re editing cells in the grid by hand until the file looks acceptable, the same thing happens as with a script nobody can read: the data gets adjusted until it satisfies rules that were never written down. The file passes, and nobody knows why.
Whichever route you take, the output is a cleaned file on somebody’s laptop. It isn’t in your system, and nothing has gone back to the client who sent the broken file.
For implementation teams, this is where most of the cost sits. The expensive part was never the 47 fixes. It was the four-day round trip:
What the client needs is a list they can work from, per row, in their own file. What the assistant produced was a paragraph in your chat window, which you now get to rewrite as an email.
All of the above happens with a person present, on one file, and it’s already harder than it should be. Recurrence adds three more problems:
A file lands in a shared mailbox at 2 a.m. on the first of the month, and a chat window doesn’t wake up. As of September 2026, Claude for Excel operates on the workbook you have open. Copilot’s Agent Mode, generally available across Word, Excel, and PowerPoint since April 2026, is explicitly a delegation model: you describe the outcome, the agent works, and you review the result. Both need somebody sitting there.
February’s file gets a fresh conversation, phrased slightly differently, and the output may differ in ways nobody notices: a rounding choice, a blank treated as zero rather than null, a date read as American rather than European.
Claude for Excel doesn’t save chat history between sessions, so whatever reasoning produced January’s clean file is gone by the time February’s arrives. For analysis, that variability stays invisible. For a monthly process where the figures must be comparable, it’s disqualifying.
When a client or an auditor asks which rows changed, under what rule, and who approved it, a chat transcript doesn’t answer. Claude for Excel currently offers no audit logs for enterprise compliance, and Microsoft’s position is that Copilot in Excel isn’t yet suited to autonomous, reproducible output without human validation.
Microsoft’s own guidance on Copilot in Excel adds an awkward detail: result quality depends heavily on how well structured the data is, and the tool struggles with unstructured files, mixed formats, inconsistent currencies, and ambiguous column names. That’s a fair description of the files an implementation team receives from its customers.
This is a fair objection, and it deserves a straight answer. Microsoft has a real automation stack – Power Automate flows, Office Scripts, scheduled triggers, Copilot agents – and ChatGPT has scheduled tasks. All of it can be pointed at files.
The question is what you’d be building: a flow that watches a folder, calls a transformation, handles the file that doesn’t match the expected structure, and routes the failure to somebody, with error handling, retries, and a log. That’s a pipeline running on good infrastructure, and it needs someone who can build and maintain it.
That brings you back to the same place. An implementation manager can’t edit an Office Script any more than she can edit a Python script. When a client renames a column, the fix becomes a change request to whoever owns the flow. When a file fails, she receives an error from a system she can’t inspect, and she still has no screen showing which rows broke which rule.
This does nothing for the single-file problem, either: automating the run doesn’t give anyone a worklist.
Client files often arrive under a data processing agreement that specifies where the data may travel. Putting a customer’s employee list or transaction extract into a general-purpose assistant may fall outside that agreement, regardless of the vendor’s own security posture. Teams typically discover this after the fact, so it’s worth raising with whoever owns your DPAs before the practice spreads across a team.
Pricing points the same way. Assistants are licensed per seat, while file processing scales with files, not people. As volume grows, you choose between buying seats for people who don’t need an assistant and routing everything through the one licensed colleague, who becomes the bottleneck.
Assistant as analyst, pipeline as operator
The teams getting the most out of these tools have stopped treating it as an either/or choice. They divide the work by what each is for.
An assistant makes an excellent analyst. It reads an unfamiliar format quickly, drafts rules well, and explains its reasoning when asked. What it can’t do is operate: hold a worklist, keep state, hand work between people, tell the customer what to fix, and do all of it the same way next month.
The mistake isn’t using AI for this work. It’s asking a strong analyst to do an operator’s job.
The first applies to a single file: could a colleague pick up where you left off?
If your progress through the errors exists only in your head and a chat window, the answer is no. That’s an interface problem rather than an intelligence problem.
The second applies to the month after: if the same file arrives again, does anything you did today still apply?
If the logic is stored, runs without you, and produces the same result, you have automation. If you’ll open a fresh conversation and work it out again, you had analysis. It’s often valuable and often the right call, but it doesn’t carry over, and next month will cost as much as this month.
See how easily you can turn your customer files into ready-to-use data with Ingestro’s AI agents.