Why ETL Tools Keep Failing at Implementation

Michael Zittermann
Michael Zittermann
Co-Founder & CEO
Last updated on
September 27, 2026
Why ETL Tools Keep Failing at Implementation

Your pipeline has run without issues for months. Same customer, same weekly file, same folder. Then, all of a sudden, a run fails, and nothing on your side changed. The customer’s system exported the file with three extra rows above the header, and the transformation looked where it always looks and found nothing there.

One data specialist at a tax compliance provider described this almost with a shrug. The flow expects the header on line five. Sometimes it arrives on line ten. Sometimes line two. The customer knows it shouldn’t move, and it moves anyway.

Inside your own systems, that’s an incident. Someone gets paged, someone ships a fix. At the customer boundary, it’s a Tuesday, and that difference is why ETL tools, iPaaS platforms, and hand-built internal pipelines keep getting bought for customer data onboarding and keep falling short.

The assumption ETL pipelines are built on

Every scheduled data pipeline rests on one premise: you know what the source looks like, and it will look the same next time.

Call it the “stable-source” assumption. It isn’t a flaw. It’s the entire design. A pipeline reads a source system, applies a transformation, writes to a target data model, and repeats on a schedule. That works because someone on your side owns the source. The schema is documented, the export is deterministic, and if it breaks, the fix happens within your organization.

Customer data onboarding rarely satisfies any of those conditions.

‍The source is unknown. An implementation manager at an HR software vendor put it plainly: there will be legacy providers his team doesn’t know about until the first customer shows up using one. You can’t design a transformation against a source system you’ve never seen.

‍The source is unstable. Format drift is the expected event, not the exception. Same customer, same export, different structure three months later.

‍Arrival isn’t a schedule. One tax compliance provider runs eight lines of business. In one, customer files arrive twice a year inside a three-month window. In another, monthly, staggered across the month to match filing deadlines. Their managed services director said something that should end the scheduling conversation: new customer onboardings can’t be foreseen or planned quarterly. A contract lands, and the work starts.

‍Nobody owns the source but your customer. The person who sent you the file doesn’t work for you. Often they can’t answer questions about their own data, because it’s been outsourced for years and the knowledge went with it. One implementation leader described a case where employees had been sending hours directly to the outgoing provider, and the customer had no idea the data existed.

An internal pipeline is a contract between two systems you own. Customer data onboarding is a contract with someone who never signed it.

The same vendor, two different problems

Here’s what makes this easy to misdiagnose. Most vendors already have a working automated data flow. It just isn’t the one that’s hurting them.

Once a customer is live, the recurring change file is standardized, narrow, and genuinely automatable. Joiners, leavers, adjustments. A thin template, a known structure, a predictable cadence. Pipelines handle that well.

The initial load is a different animal. It’s wide, it happens once, and it’s formatted entirely by whatever the customer’s legacy system produced on its way out the door. An implementation lead at a global provider described the onboarding template as carrying far more fields than the steady-state one, then noted that once they’re live, it all gets easier.

Two of these deserve emphasis together, because the same pattern showed up at a different kind of company. A product manager at an HR software vendor explained that fixed monthly amounts recur automatically without anyone touching them, while variable inputs coming from a third-party time system are exported by a person and re-imported by hand. Same product, same customer, same month. The automation holds where the data is internal and stable, and it collapses where the data comes from a system nobody at the vendor controls.

The pipeline model isn’t bad. Implementation is simply the phase where none of its preconditions hold.

What it costs to build on the assumption anyway

Teams build on it anyway, because it’s the only kind of tooling most of the market sells. Four costs follow, roughly in order of visibility.

The mappings multiply, and customers aren’t the multiplier

The real multiplier is customers times legacy systems times countries times service partners times target templates. One enterprise moving off an incumbent ETL tool counted something on the order of six hundred transformation flows already in production. An HR vendor told us they maintain roughly a hundred import templates, of which ten to fifteen get used on any single implementation. A global provider keeps separate local information workbooks that go all the way down to country and, where they work with more than one partner in a market, down to partner.

The obvious response is to build one connector per legacy system and stop rebuilding. That doesn’t survive contact either. An implementation leader described two customers running the same legacy software whose exports still differ, field by field, because each one configured it their own way.

The rules end up in places nobody audits

A jurisdiction encoded in a filename. A required worksheet name. A folder per customer on a shared drive, swept twice a day by a server. Each of those is a condition imposed on a file your organization doesn’t produce, enforced by a convention that someone at another company either follows or doesn’t.

The logic stops being readable

Rules accumulate for years inside a tool where changing one requires knowing how it was built in the first place, which typically means knowing who built it. One team’s institutional knowledge existed only as machine-readable pipeline definitions. To explain to anyone what those flows did, they had to export the definitions and run them through a language model. Elsewhere, the logic was never written down at all. A product manager at a global employment provider described market expertise living in the head of the coordinator who covers that market, and what happens the week that person is on leave.

The bill arrives on the services P&L, not the engineering backlog

This is why the problem can stay invisible to the people who would fund fixing it. It doesn’t surface as a ticket queue. It surfaces as implementation consultants spending hours per customer reformatting spreadsheets. At vendors who bundle implementation into the subscription price, and many do, each of those hours comes straight out of margin. Several services leaders described the same thing independently: the data handoff is the one part of the project that slows go-live, and the rest of it runs fine.

Three fixes that look obvious from a distance

Each of these is reasonable. Each has been tried by someone in our customer conversations. It’s worth being precise about how far each one gets.

Mandate the format

This works, partially, and then stops. One vendor did it properly: they banned PDFs and narrowed to three accepted file types. Variation dropped and didn’t disappear. The ban turned out to be size-dependent, because small customers still got exceptions. Months later, the banned format was the case they most wanted solved, since it came from the one legacy provider causing their team the most pain.

Another had gone further and written format requirements into new contracts. It still didn’t reach the tail. Their principal engineer explained why: some counterparties employ a handful of people in a single country and can’t produce anything other than what their system gives them. A format mandate doesn’t close the distance between your target structure and your customer’s export. It moves where the mess sits.

There’s one real exception, and it proves the point. Where a regulator defines the schema, the source genuinely is stable. Stability comes from authority over the source, and you have authority over your own systems and none at all over your customers’.

Route it through the API

Several vendors we spoke with have an import API and don’t use it for onboarding. The reason is structural: an API presupposes a caller who already knows your target structure and can produce it. The party holding the data is typically the incumbent provider being replaced, and they have no obligation to help. One implementation leader noted that customers are sometimes reluctant to even request a full extract, because the request itself signals to their current provider that the relationship is ending.

Validate on the way in

The most sophisticated version of the wrong answer. One HR software vendor built an Excel validation plugin years ago, and it works exactly as designed. Their senior product manager identified the problem with unusual clarity: by the time the data reaches the point of validation, a person has already cleaned it up by hand, so none of that work is saved. Validation sits downstream of the expensive part.

These aren’t three failed fixes. They are three attempts to make a source you don’t own behave like a source you do.

What the work asks for instead

Strip out the tooling assumptions, and the real requirements look different from what a pipeline product is designed to deliver.

Failure has to be a normal state with a first-class repair path, rather than an exception handler. When a customer file breaks a run, someone has to decide whether the change is a one-off or the new normal from that customer. A scheduled job can’t make that call. A person talking to the customer can.

The loop runs through someone outside your company. Validation that finds a real problem often has to go back to the customer for clarification, and then wait. A transformation graph has no step for that.

Changes need consent and a record. You’re reformatting data on someone else’s behalf, and one implementation leader was firm about it: if fifteen values were corrected, the customer needs to see the fifteen and approve them. Nobody asks permission to reformat their own warehouse table. At the customer boundary, it’s the baseline.

The people doing this work aren’t engineers. At the vendors we spoke with, it sits with implementation consultants, onboarding managers, and operations leads. One principal engineer made it a buying criterion, asking whether his own team would be needed at all and hoping the answer was no.

And the unit of reuse has to change. A connector is the wrong unit, because there’s nothing stable to connect to. What needs to accumulate is recognition of formats you’ve already seen.

Where AI data workflows fit the work in front of you

That last requirement is the one that reframes everything, and it’s where AI data workflows earn their place.

An ETL flow is a fixed instruction set: take this column, apply this rule, write this field. It’s precise, it’s auditable, and it’s brittle by design, because every instruction encodes an assumption about a source you don’t control.

An AI data workflow changes the order of operations. AI data mapping proposes the match between source fields and your target data model, reading headers, sample values, and the formats it has processed before.

Your team reviews what’s proposed, and deterministic, inspectable rules carry out the transformation, with validation and AI-powered cleaning applied consistently on every run. The judgment adapts. The execution doesn’t, and nothing is applied to customer data before someone confirms it.

What that gives you in practice: a file arriving in a structure nobody anticipated becomes a normal input rather than a failed run. Recognition carries forward across customers instead of dying within a single flow. Your operations leads can adjust a mapping without opening a ticket. Every row-level change is logged, so when a customer asks what happened to a record, you have an answer – which is more than a manual spreadsheet process can offer.

That’s the ground Ingestro is built on, and it’s worth being straight about the boundaries. Merging several files without a shared identifier remains hard, and sometimes the honest answer to a customer is that it can’t be done reliably. Genuinely unstructured input still needs human judgment. A small amount of logic still gets written as code. Per-customer work doesn’t disappear.

The argument was never that it disappears. It’s that the schedule-and-transform model has nowhere to put per-customer work, so it hides in consultant hours and engineering tickets where nobody counts it.

Questions worth putting to any vendor, including us

  • What happens with a file structure the system has never encountered? Not throughput, but behavior with the unknown.
  • When a run fails on a bad file, who gets notified, what do they see, and can they repair that run without touching the configuration?
  • Is a defect affecting an entire column fixable in one action, or row by row?
  • Can an operations person build and change a mapping without a developer?
  • Does mapping work transfer between the manual upload path and the automated one, or is it built twice?
  • What’s the audit record when a value is changed on a customer’s behalf?
  • Can it run inside your own infrastructure, and can AI processing be switched off for a specific customer who prohibits it?
  • What’s the real setup time per source, measured on your files rather than a demo file?

A short audit you can run this week

  • Of your last ten implementations, how many reused a mapping built for a previous customer without modification?
  • When a customer file breaks a run, how long until someone at that customer gets asked about it?
  • Can you name who owns the transformation logic for your largest account, and can anyone else read it?
  • Is the data handoff on your implementation plan as a step with an owner and a duration, or is it an assumption?

Takeaway

The takeaway for leadership is simple. The question was never which ETL tool to standardize on. It’s whether the phase your revenue timing depends on is running on a model built for sources you own.

Implementation is the one part of your business where the stable-source assumption was never true. Everything that follows from pretending otherwise is a cost your services team is already paying, quietly, on every project.

Faster and more secure customer data onboarding
Turn complex customer files across sources and formats into clean data flows through AI automation.
Explore solutions

See how easily you can turn your customer files into ready-to-use data with Ingestro’s AI agents.

Keep exploring

icon