Product & Engineering

Product & Engineering

Garbage In, Confidently Wrong Out: Why Context Quality Matters More Than Context Volume

Garbage In, Confidently Wrong Out: Why Context Quality Matters More Than Context Volume

Garbage In, Confidently Wrong Out: Why Context Quality Matters More Than Context Volume

Rohan Mahajan

Rohan Mahajan

7 Min

7 Min

Build Connected Systems with Tartan

Automate workflows with integrated data across your customer applications at scale

The pitch behind almost every “connect everything” integration tool sounds reasonable on the surface: give your AI system access to more of your data, and it will reason better. More systems connected, more context available, better outputs. It is an easy story to sell, because it maps onto an intuition most people already have, that more information should produce better decisions.

Inside a real enterprise, that intuition breaks. Most enterprise data is not a clean, single source of truth waiting to be connected. It is stale in some systems, duplicated across others, and quietly contradictory between the two. Feeding all of it to a model does not make the model smarter. It gives the model more raw material to be confidently wrong with, and confidently wrong is a much harder failure mode to catch than a system that simply does not know.

The firehose assumption

Firehose integrations rest on an assumption that does not hold up inside most enterprises: that every system of record is accurate, current, and consistent with every other system it might be cross-referenced against. In practice, none of that is reliably true.

A few patterns show up constantly once you actually look:

  • Staleness. An employee’s status was updated in the HRMS three days ago. The CRM record referencing that employee has not synced since last quarter. Both are “real” data. Only one reflects reality right now.

  • Duplication. The same customer exists as two or three separate records across a CRM, a core banking system, and a KYC vendor’s database, each with slightly different details, none of them flagged as duplicates of each other.

  • Contradiction. One system says a loan is active. Another, updated on a different cycle, still shows it as pending. Neither system is technically wrong. They are just not synchronized with each other, and nothing upstream resolves that conflict before a model sees both.

A model handed raw access to all of this is not reasoning over “more context.” It is reasoning over multiple, sometimes conflicting, versions of the truth, with no signal about which one is current, which one is authoritative, or whether any of them should be trusted at all.

Why this fails silently

This is the part that makes the firehose approach genuinely dangerous rather than just inefficient. A model given contradictory or stale inputs rarely responds by flagging the contradiction. It picks one version, often whichever surfaces first or fits the prompt best, and produces a fluent, confident answer built on top of it.

Nothing about that output looks uncertain. It reads exactly like an answer built on clean, verified data. The failure is invisible until someone downstream acts on it and the decision turns out to be wrong, at which point tracing it back to “the model picked the stale record instead of the current one” is its own investigation, often happening well after the decision already had consequences.

This is a materially worse failure mode than a system that simply lacks data. A model that says “I don’t have enough information” is at least honest about its own limits. A model that confidently answers using the wrong version of a contradictory record gives no such warning, and in a regulated workflow, that confidence is exactly what makes the output dangerous to trust.

The volume-quality tradeoff nobody prices in

Most conversations about giving AI systems context focus entirely on breadth: how many systems are connected, how much data is reachable. Almost none of them ask a more basic question: of the data being connected, how much of it is actually current, deduplicated, and verified against a source of truth.

These are different axes, and they do not move together. It is entirely possible, common even, to have an integration that connects to ten systems and produces worse outputs than one that connects to three, if those three are clean and the ten are not. Breadth without quality does not average out to something reasonable. It just means more of the model’s reasoning is now resting on data nobody has verified.

The uncomfortable tradeoff most vendors do not want to say out loud: getting to clean, verified, deduplicated context across even a handful of systems is genuinely hard, slower, and more expensive than wiring up a broad connector and calling it context. Selling breadth is easy. Selling curation is a harder pitch, because it sounds like doing less, even though it is the version that actually produces trustworthy outputs.

What curated context actually looks like

A smaller, curated context set is not simply “less data.” It is a specific set of properties that raw access does not automatically provide:

  • A resolved source of truth per fact. When the same fact exists in multiple systems, curated context means deciding, in advance, which system is authoritative for that fact, rather than leaving the model to guess between conflicting versions at inference time.

  • Deduplication before retrieval, not after. Duplicate customer or employee records get resolved into a single canonical entity before a model ever reasons over them, not flagged as a downstream cleanup task once something has already gone wrong.

  • Freshness as a first-class property, not an afterthought. Data that has gone stale gets marked or excluded, rather than treated as equally valid to a record updated an hour ago.

  • Verification, not just access. A field being reachable through an API is not the same as that field having been checked against its source of truth. Curated context distinguishes between “we can technically retrieve this” and “we have confirmed this is accurate right now.”

None of this is a smaller version of the connect-everything approach. It is a different kind of work entirely, closer to building a translation and normalization layer than wiring up a connector, and it is the difference between a model that reasons over something trustworthy and one that reasons over whatever happened to be reachable.

Why this matters more as stakes go up

In a low-stakes setting, a model occasionally picking the wrong version of a fact is a minor annoyance. In BFSI, it is the difference between a correct lending decision and an incorrect one, a valid KYC clearance and a false one. The cost of a confidently wrong answer scales directly with how much weight gets put on the decision it produces, and lending, underwriting, and claims decisions carry about as much weight as enterprise data gets asked to support.

This is exactly why the firehose approach is the most tempting in precisely the environments where it is most dangerous. Regulated industries have the most systems, the most legacy inconsistency, and the most pressure to move fast on AI adoption, which makes broad, quick integrations attractive and makes the consequences of skipping curation the most severe.

The contrarian case, stated plainly

Most vendors selling “give your AI more context” are, deliberately or not, selling volume, because volume is easier to demo and easier to measure. Connected systems, integrations live, data flowing. None of that tells you whether what is flowing is accurate, current, or free of contradiction.

The better question to ask of any context strategy is not how much is connected. It is how much of what is connected has actually been verified, deduplicated, and resolved to a single source of truth. A smaller set of genuinely clean, current, verified data will consistently outperform a larger set of unverified, potentially contradictory data, because the model reasoning over the smaller set is working from something real, and the model reasoning over the larger set is working from a plausible-sounding average of several things that might all be partially wrong.

Garbage in, confidently wrong out is not a rare edge case. It is the default outcome of connecting first and verifying never, and it is exactly why context quality, not context volume, is the variable that actually determines whether an AI system can be trusted with a real decision.

One platform. Across workflows.

One platform. Across workflows.

Tartan helps teams integrate, enrich, and validate critical customer data across workflows, not as a one-off step but as an infrastructure layer.

Tartan helps teams integrate, enrich, and validate critical customer data across workflows, not as a one-off step but as an infrastructure layer.

Tartan helps teams integrate, enrich, and validate critical customer data across workflows, not as a one-off step but as an infrastructure layer.