# When Your AI Agent Can’t See What It Breaks
Modern code-generation tools have made impressive leaps. Ask one to write a function and it delivers. Ask it what downstream effects a modification will have, and it improvises — often confidently and often incorrectly.
This gap between generation and impact analysis is not a mystery. It is a structural problem rooted in how we organize software today, and it is silently eroding the trust developers place in AI-assisted workflows.
## The Two Questions Code Tools Answer
Every code-retrieval system currently deployed answers one question: “Given this description, which parts of the codebase look most like what was described?” The engine embeds code chunks, embeds the query, finds the nearest neighbours, reranks, widens the window. Whatever improvements get layered on top, they all serve the same goal: returning a set of code that resembles what you asked about.
That is the wrong question when you are trying to change something safely. What you actually need is a different set entirely: the set of code that will stop working when the thing you are changing stops working.
Those two sets overlap, but the overlap is almost always smaller than engineering teams assume.
## Where the Blind Spots Hide
Consider a component library published as a package and consumed by multiple applications. The library’s repository contains no references to the apps that pull it in. An application’s repository contains the package name in its manifest, and perhaps a string reference in a configuration file — but no indexer looking at the app’s source code will ever find the library’s internal structure, because it lives behind a dependency boundary that resolves into a local cache rather than a readable source tree.
This is not a flaw in any particular editor or assistant. It is what happens when well-architected systems use explicit contracts, interfaces, and abstraction layers. The more disciplined the design, the less vocabulary overlap exists between the thing being changed and the things that depend on it. A shared type definition and the component that consumes it may have no textual similarity at all, yet changing the type’s shape breaks the component. A similarity index cannot reason about that relationship because the relationship lives in the system, not in the text.
## The Hidden Architecture of Composition
The problem deepens when systems stop importing units directly and start composing them at runtime. A platform might register frontend applications, backend services, and gateways by identity rather than by value — using string references in configuration arrays, dynamic import resolution, or environment-level wiring. There is no import statement, no type signature crossing the boundary, and no “find all references” that can trace it. The connection is real, it is enforced by the deployed system, and it is invisible to every static analysis tool that looks for declarations.
This kind of wiring is not accidental. It is the natural result of independently deployable units that agree on a protocol but do not share a compilation boundary. The contract between them lives in a runtime agreement — an HTTP response shape, a message queue schema, a slot registration convention — and that agreement is often recorded nowhere an indexer can reach.
## What a Dependency Record Actually Provides
The alternative to runtime discovery is recording dependencies at the moment a unit is built and published. Each version of a component captures which other components it relies on, and each consumer’s record captures that same relationship from the other direction. When an agent needs to know what breaks if it modifies a piece of shared infrastructure, it does not need to guess, it needs to look up the recorded edge.
This is fundamentally different from what most existing tools do. Language servers resolve references inside a single repository. Build tools like Nx and Turborepo maintain project graphs from TypeScript imports. Cross-repository tools such as Sourcegraph and GitHub’s “Used by” feature derive their graphs from whatever happens to be indexed in each direction. The critical limitation is that these tools analyze text at read time and require both sides of a relationship to be visible simultaneously.
Recording edges when versions are created sidesteps both requirements. It does not require analyzing source code at query time, and it does not require both the producer and the consumer to be accessible to the same indexing process. The data is a fact, not a derivation.
## Why Team Boundaries Make This Worse
In a solo project, publication boundaries are deliberate choices, and a single developer can keep the graph manageable. In an organization with multiple teams, every team boundary becomes a publication boundary by default. The number of cross-boundary relationships stops reflecting engineering discipline and starts reflecting the organizational chart.
When a shared contract changes, the person making that change has almost certainly never seen the majority of the code that holds the other end. Visibility, not intent, is the constraint. A queryable graph transforms “nobody told me this would break” into a lookup that surfaces the blast radius automatically.
## What This Does Not Solve
A dependency graph tells you what is affected by a change. It does not tell you whether the change itself is correct, safe, or desirable. Impact analysis and correctness analysis are different concerns, and neither one replaces the other.
It also cannot help in codebases that lack clear component boundaries in the first place. If a system has no modular seams, there is no graph to query, and no recorded edge will expose a contract that was never declared.
Finally, this is a perspective from one system’s experience. The patterns described here are general, but the specific implementation details are particular to the architecture being discussed. The value is in the shape of the insight, not in any single tool’s claims.
## The Core Problem Remains
Code generation has solved the “write it from scratch” problem remarkably well. The “change something without breaking the rest of the system” problem remains unsolved, and the reason is not model capability. It is that agents are being asked to answer a system-structure question using a text-similarity engine, and the answer they produce reflects what is visible in the immediate context — which is never the whole system.
The fix is not to make similarity search better. The fix is to ask a different kind of question and to build the infrastructure that lets agents ask it. When edges are recorded at build time and stored in a form that survives repository boundaries, an agent can look up impact instead of guessing at it. That does not eliminate the need for human judgment, but it does eliminate an entire class of false confidence that currently makes developers trust the output less, not more.
Structuring your agent’s queries to use dependency information rather than textual similarity is what makes the difference. A system the agent cannot query does not limit the agent. It limits the developer who built the system.
## Frequently Asked Questions
**Why can’t current AI assistants figure out what breaks when I make a change?**
They rely on similarity-based retrieval, which finds code that looks like your change. The code that actually breaks is often textually unrelated to what you are modifying — for example, a consumer that depends on an interface you changed but never mentions your implementation by name.
**Doesn’t every project already have an import graph?**
It has an import graph within its own repository. Import graphs do not cross package boundaries, private repositories, or organization borders. They also cannot capture edges formed by runtime composition, configuration, or deploy-time wiring.
**What is the difference between a similarity index and a dependency graph?**
A similarity index answers “what looks like this?” A dependency graph answers “what depends on this?” Those are different questions, and the one that matters for safe modification is almost never the first one.
**Are tools like Nx or Sourcegraph solving this already?**
They solve part of it — specifically, the part that lives inside a single repository or package boundary. They fail at cross-repo and cross-organization visibility, and they derive their graphs at read time rather than recording them at build time.
**Does this approach require a specific platform or product?**
The pattern does not. It requires that dependency edges are recorded somewhere when units are composed and published. Any system that does this at a granularity fine enough to expose component-level relationships can power the same kinds of queries.
**Why does modularity make this harder instead of easier?**
Modularity replaces direct imports with interfaces, contracts, and abstraction layers. Those are excellent for maintainability, but they deliberately remove the textual overlap that similarity-based retrieval relies on. The better the design, the more invisible the dependencies become to a text-only index.
## Conclusion
The next frontier of AI-assisted development is not better code generation. It is better impact awareness. The tools we already have can generate code impressively; what is missing is a way for those tools to understand the structure they are operating inside.
Recording dependency edges at build time, making them queryable across team and repository boundaries, and teaching agents to look up impact rather than guess at it are the three pieces of a solution that is simpler than most people assume. The hardest part is not the technology — it is the decision to treat structural relationships as first-class data rather than something to be rediscovered from text every time a question is asked.
The developers who adopt this pattern will see fewer surprise breakages, more confidence in their tooling, and a clearer line between what the machine can verify and what it still needs a human to judge. That is a gap well worth closing.
Thank you for reading



