Software budgets are shrinking. Nearly $1 trillion in market value came off software and services stocks in the first weeks of February 2026, and the average enterprise now runs fewer SaaS applications than it did two years ago. One line on every data team's invoice didn't get the memo: the pipes that move data into the warehouse.


The Narrative Doesn't Match the Invoice

The SaaSpocalypse story is real, and it is about seats. Per-user productivity tools, CRMs, and workflow software are the categories bleeding value as AI agents take over tasks that used to require a named license. TechCrunch, citing Reuters, put the February 2026 selloff at nearly $1 trillion in market value wiped from software and services stocks. Zylo's 2026 SaaS Management Index shows the average enterprise's app count finally shrinking after years of sprawl.

Data ingestion tooling is running the opposite direction. Gartner VP Mike Tucciarone told CIO.com that SaaS subscription costs from large vendors rose 10% to 20% this year, against IT budget growth projections of just 2.8%. That gap alone should worry a CFO. But the number underneath it is worse: Unravel Data CEO Kunal Agarwal put the rise in cloud data warehouse, lakehouse, and analytics platform costs at 30% to 50% in the past year - two to five times the pace of SaaS overall.

These are not the same market moving in two directions by coincidence. Per-seat software is priced by headcount, and AI is cutting headcount for the tasks those tools supported. Data infrastructure is priced by volume, and AI is doing the opposite to volume: every agent that reads a CRM record, writes a log line, or triggers a sync adds rows, not seats. Cutting a Salesforce license does nothing to a bill that scales with data moved.

Gartner (Tucciarone, via CIO.com, 2026) and Unravel Data (Agarwal, via CIO.com, 2026) - SaaS and IT budget figures from the same CIO.com report.

Per-seat SaaS is a headcount tax. Data ingestion tooling is a volume tax. AI cuts headcount for specific tasks almost immediately; it multiplies data volume even faster - which is why the bill that's actually growing is the one nobody budgeted for.


What Fivetran Actually Changed in March 2025

Fivetran's pricing runs on Monthly Active Rows: distinct primary keys touched by an insert, update, or delete in a given month, counted once no matter how many times that row syncs. Until March 2025, MAR was pooled at the account level - five connectors sharing one volume curve, with bulk-pricing discounts kicking in as the combined total grew.

Since March 2025, Fivetran's own documentation confirms MAR is tracked separately for each connection. A team running ten moderate-volume connectors no longer gets one discount curve for the combined total; it gets ten separate curves, none of which individually reaches the volume where the price per row drops. Fivetran also added a $5 base charge per connection between 1 and 1 million MAR, and - as of a further January 2026 change - deletes now count as billable MAR alongside inserts and updates, where they previously didn't always.

None of this shows up as a single headline percentage, and it would be dishonest to invent one. What Weld's own pricing breakdown found instead is qualitative but blunt: the change "increased costs significantly for teams running many connectors with moderate volumes each" - precisely the shape of a mid-sized data team's stack, not the enterprise accounts with one dominant, high-volume source.

Run the mechanism forward with round numbers. A team with 8 connectors, each generating enough activity to owe the $5 base charge plus a modest per-row fee, now pays that base charge eight times over before a single row is priced - a fixed cost that didn't exist as a per-connection line item under the old pooled model. Multiply that across a data stack that keeps adding SaaS-source connectors (Salesforce, Stripe, HubSpot, ad platforms, support tools) and the base charges alone start to look like a second subscription fee, layered under the usage fee.


The Merger That Explains the Pricing Confidence

In October 2025, Fivetran and dbt Labs announced an all-stock merger, closing in June 2026, forming a company approaching $600 million in combined annual recurring revenue. Fivetran's own press release frames it as building "the data infrastructure for trusted AI agents." Sacra puts Fivetran's last private valuation at $5.6 billion, a 59x revenue multiple that only makes sense if the market believes pricing power, not customer growth, drives the next stage.

The two products were never technically identical, but they were priced by two separate vendors that a customer could negotiate independently: Fivetran moved data in, dbt transformed it. Buying both from separate companies created a natural price check - a team unhappy with one vendor's roadmap or pricing could swap it without touching the other. A combined ingestion-plus-transformation vendor removes that check. It doesn't need to prove per-tool value against a competitor; it needs to prove stack value against the cost of switching two tools at once.

This is the same consolidation logic Fivetran already applied internally in March 2025 when it collapsed account-level pooling into connector-level billing: reduce the customer's ability to average costs down across their own usage, and reduce it again across the vendor relationship itself.


Stitch's Quiet Wind-Down Is the Warning Most Teams Missed

Qlik acquired Talend, and Talend had already absorbed Stitch Data years earlier. Qlik's own migration center documentation is now actively walking Stitch customers through a formal transition to Qlik Talend Cloud - inventory tools, schema-transition guides, a dedicated migration toolkit. Qlik has not issued a blunt "Stitch is dead" announcement, and it would be inaccurate to claim one exists. What the documentation shows is a vendor-driven migration in progress, not a customer-initiated one.

That distinction matters more than a formal sunset date would. A customer-initiated migration happens on the customer's timeline, driven by a genuine need. A vendor-driven migration happens on the vendor's roadmap, and the customer absorbs the engineering hours whether or not the current setup was working fine. Teams that built ELT pipelines on Stitch because it was the simple, cheap option in 2018 are now spending real migration time on a transition they didn't request, for a product that was supposed to be the low-maintenance choice in the first place.

The lesson generalizes past Stitch. In a market consolidating around fewer, larger data infrastructure vendors, the "safe, boring, cheap" tool is exactly the one most likely to get folded into someone else's platform on someone else's timeline.


The Build-vs-Buy Line Just Moved

Airbyte's open-source core is genuinely free to run, and its cloud pricing shifted too - from a credit-based model to capacity-based pricing in February 2025, a change explicitly framed by Airbyte as removing "usage-based surprises." That framing only makes sense as a response to exactly the kind of unpredictable billing Fivetran customers were describing at the same time.

But "free" and "self-hosted" are not the same word. Integrate.io's own cost breakdown puts a production Kubernetes deployment of Airbyte at $500 to $3,000+ per month in infrastructure alone, plus 20 to 40 hours per month of engineering maintenance. Priced against a loaded data engineer cost - Indeed, ZipRecruiter, and Glassdoor all put average total compensation in the $130,000-$137,000 range, which lands near $85 an hour once benefits and overhead are counted - that maintenance time alone runs roughly $1,700 to $3,400 a month. Add the infrastructure spend and the honest total cost of "free" self-hosted ingestion is closer to $2,200 to $6,400 a month, not zero.

Integrate.io (2026) infrastructure estimate; engineering time costed at ~$85/hr loaded, derived from Indeed/ZipRecruiter/Glassdoor 2026 data engineer salary data.

That range is the real comparison point against a managed vendor's invoice - not zero. The math only favors self-hosting once a team's Fivetran-style bill, inflated by per-connector base charges and lost pooling discounts, climbs past that $2,200-$6,400 floor. For a handful of connectors at low volume, managed still wins easily. For the mid-sized stack with eight or more moderate-volume connectors - exactly the shape March 2025's repricing hit hardest - the crossover arrives faster than most teams have modeled.

no

yes

no

yes

8+ connectors,
moderate volume each?

Stay managed
Fivetran/Airbyte Cloud

Already running an
orchestrator in-house?

Model the crossover
before switching

Self-host Airbyte
marginal orchestration cost

The crossover assumes infrastructure and engineering-time costs from the chart above; teams already running Airflow, Prefect, or Dagster absorb ingestion as marginal load on existing orchestration rather than a new maintenance burden.

That second question is the one this article can't answer in isolation. A team with no in-house orchestrator is comparing Fivetran's invoice against Airbyte's infrastructure cost plus the cost of building scheduling, retries, and monitoring from nothing. A team that already runs Airflow, Prefect, or Dagster is comparing it against marginal load on a system that already exists. The orchestrator decision and the ingestion decision aren't separate purchases - they're one build-vs-buy calculation split across two vendor relationships.


Who the Crossover Doesn't Work For

That decision tree assumes something most companies don't have: spare engineering capacity. Carta's H2 2025 compensation data puts the median seed-stage startup at just four employees, with engineers accounting for roughly 30% of first-half-2025 hiring - call it one, maybe two people, both fully committed to product. Series A headcount has actually shrunk, from a median of 57 employees in 2020 to about 47 in 2025, as funding tightened and teams stayed leaner for longer. For a company at that stage, the $2,200-$6,400 monthly self-hosting cost from the chart above isn't the real comparison. The real comparison is a $130,000+ hire against a Fivetran invoice - and self-hosting isn't a line item, it's a headcount decision most funded companies aren't in a position to make.

The obvious counterargument is that AI coding agents have collapsed the effort side of that equation - a generalist engineer with an agentic coding assistant can now stand up and maintain infrastructure that used to require a dedicated platform hire. That's true, but only under a condition the 2025 DORA report states plainly: "AI's primary role is as an amplifier, magnifying an organization's existing strengths and weaknesses." Teams with mature platforms, standardized environments, and established deployment pipelines convert AI into genuine toil reduction. InfoQ's review of the same report describes the downside case in blunter terms: for teams with fragmented tooling or unclear process, AI can accelerate the creation of technical debt, increase code review complexity, and introduce instability into systems that were already fragile.

The teams AI actually helps run self-hosted infrastructure well are the ones that already had the platform maturity to negotiate a better managed contract in the first place. The four-person seed startup weighing a hire against an invoice is not that team.

This is the same failure mode covered in the hidden cost of AI-generated code: locally correct output that looks like it removed the need for expertise, while quietly shifting the maintenance burden to whoever inherits it later. An AI agent can generate a working Airbyte deployment manifest in an afternoon. Whether that manifest survives a schema change, a credential rotation, or a connector failure at 2 a.m. depends on exactly the operational maturity DORA says AI doesn't create on its own - it just makes the absence of that maturity less visible until it isn't.

None of this means self-hosting is a bad option. It means the crossover math from the previous section is necessary but not sufficient. A team below roughly 15-20 engineers, without an existing platform function, should read the self-hosted cost range as a floor that assumes competence it doesn't yet have - not a ceiling it can budget around.


Why This Keeps Happening

The data integration market is not shrinking alongside per-seat SaaS. Research and Markets puts it at $15.13 billion in 2025, growing to $17.18 billion in 2026 - a 13.5% CAGR, in the same window the broader software market lost nearly a trillion dollars in valuation. That is not two markets on the same cycle; it is one market being redefined by what AI actually consumes.

Research and Markets, Data Integration Market Report (2026).

Weld's own pricing review surfaced a customer describing "a huge spike (more than double)" in their bill, attributed to the 2025-2026 pricing changes collectively - a single account, not a controlled study, but a real number from a real invoice rather than a modeled estimate. That's the pattern the connector-level math predicts: every new connector is now priced as if it were the only one, and a team that keeps adding sources without re-checking the base-charge math will find that out from the invoice, not the pricing page.

Unravel Data's Kunal Agarwal estimates 20% to 40% of data infrastructure spend is simply waste - unused connectors, redundant syncs, tables replicated but never queried. That waste is easier to justify when the bill is one pooled number growing slowly. It gets much harder to ignore once every connector carries its own base charge and its own pricing curve, visible on its own line - a discipline the same cost-control instinct behind cutting cloud spend without freezing delivery applies just as directly to the data layer, and one that only pays off once the pipeline itself is producing numbers worth trusting, the actual subject of building a SaaS metrics stack you can defend.

The SaaSpocalypse is cutting the tools that were never the real cost. Per-seat licenses are easy to cancel because a person can stop logging in. A data pipeline can't stop moving rows without breaking the reporting a company still depends on - and that's exactly the leverage a consolidating, usage-priced vendor doesn't have to give back.


Sources

  1. Fivetran - Usage-Based Pricing / Monthly Active Rows (2026) - primary documentation on per-connection MAR, the $5 base charge, and how inserts/updates/deletes are counted
  2. Weld - Fivetran Pricing Explained (2026) - breakdown of the account-to-connector MAR shift and a customer-reported cost spike
  3. Rivery - Fivetran Is Changing Their Pricing Model (2025) - context on data source sprawl driving exposure to the March 2025 change
  4. Fivetran - Fivetran + dbt Labs Complete Merger press release (2025/2026) - primary source on the merger, combined ~$600M ARR, and strategic framing
  5. Sacra - Fivetran revenue, valuation & funding profile - $5.6B last private valuation at a 59x revenue multiple
  6. Qlik - Preparing to Migrate (Stitch to Qlik Talend Cloud) - primary documentation on the vendor-driven Stitch migration
  7. Integrate.io - Airbyte Pricing: How Much Does Airbyte Really Cost in 2026 - self-hosted infrastructure and engineering-maintenance cost estimates
  8. Hevo Data - Airbyte Pricing in 2026 - detail on the February 2025 shift from credit-based to capacity-based pricing
  9. TechCrunch - SaaS In, SaaS Out: What's Driving the SaaSpocalypse (2026) - the ~$1 trillion market value figure, citing Reuters
  10. CIO.com - SaaS Price Hikes Put CIOs' Budgets in a Bind (2026) - Gartner VP Mike Tucciarone and Unravel Data CEO Kunal Agarwal on SaaS vs. data infrastructure cost growth
  11. Research and Markets - Data Integration Market Report (2026) - $15.13B (2025) to $17.18B (2026), 13.5% CAGR
  12. SaaS Mag - The Great SaaS Rebundling (citing Zylo's 2026 SaaS Management Index) - average enterprise app count shrinking
  13. Indeed - Data Engineer Salary in the United States (2026) - baseline for loaded engineering-hour cost calculation
  14. Carta - State of Startup Compensation: H2 2025 - median seed-stage team size and H1 2025 engineering hiring share
  15. DORA - State of AI-Assisted Software Development 2025 - AI as an amplifier of existing organizational strengths and weaknesses

Working through the challenges in this post? I help engineering leaders and CTOs navigate complex technical decisions and scale high-performing teams. Schedule a consultation →