Claude Opus 5 can keep far more code and task history in working context than a small-window model, but that does not make an unfamiliar repository self-explanatory. For a consequential multi-file change, give the model a bounded architecture map, the dependency and test evidence for the change, and a human review boundary. Evaluate the resulting plan and diff, not the size of the prompt.

The Short Answer: A Better Model Still Needs Repository Context

Anthropic released Claude Opus 5 on July 24, 2026. The current API identifier is claude-opus-5; Anthropic documents a one-million-token context window and a 128K maximum output for that model. The release date is useful for the related query “Claude Opus 5 release date,” but it is not evidence that the model understands every codebase.

A context window is working memory for one request or conversation. Anthropic notes that it includes system instructions, messages, tool definitions, tool results, documents, and generated output. It is not a durable, verified model of a repository. More tokens can make a larger change visible; they cannot say which dependency is authoritative, whether a migration was superseded, or which test is the release gate.

Treat the model as an investigator with a large desk. Put the current evidence on that desk, but retain the source locations and revision that make the evidence checkable.

Follow One Change Across a Large Repository

Consider a request to add a field to a customer profile returned by an API. The agent may see the route handler, a schema, and a UI component. That is enough to propose an edit, but it may omit a generated client, an authorization policy, a backfill, a contract test, or a consumer in another service.

The files visible to the model

Visible material helps the agent form a hypothesis: the ticket, affected files, selected symbols, a module map, and a small set of tests. A strong model can compare these artifacts, call tools, and ask for missing evidence. Claude’s API supports caller-defined and Anthropic-hosted tools, so an implementation can retrieve exact repository facts instead of relying only on the initial prompt.

The dependencies hidden outside the prompt

The hidden set is where codebase work fails. It includes callers outside the current folder, generated artifacts, deployment assumptions, ownership rules, and decisions encoded in reviews or design records. Long context reduces the need to compress all of this into a few snippets. It does not rank the facts by authority or guarantee that an old decision still applies.

Editorial diagram of a bounded change request moving from visible files to dependency and test evidence

Four Levels of Codebase Understanding

Use four levels to distinguish plausible text completion from a reviewable repository conclusion.

Reading individual files

The first level is local: identify a function, type, configuration key, or test assertion. It is useful for a contained bug fix, but a correct local edit can still violate a contract elsewhere.

Navigating modules and services

The second level connects imports, calls, ownership, and runtime boundaries. The question changes from “what does this function do?” to “which modules consume this behavior, and where is the contract enforced?” A repository map and dependency evidence belong here.

Recovering design intent

The third level adds decisions: why an interface is versioned, why a fallback exists, or why a package is deliberately not shared. Store such records with a source, date, owner, and revision. Do not turn an old chat summary into an unqualified fact.

Predicting change impact

The fourth level combines current code, contracts, tests, and decisions into a change plan. It is the hardest level because it needs negative evidence too: a search result that found no consumer is not proof that no consumer exists. Ask the agent to list its evidence and its uncertainty before editing.

For the operational package behind these levels, see Claude Opus context engineering.

Where Long Context Helps—and Where It Breaks Down

Long context helps when the task needs several coherent artifacts at once: a design note, the relevant API, tests, recent diffs, and tool output. It can reduce repeated copying between turns and lets the model compare distant files without a premature summary.

But context capacity is not a relevance ranking. Anthropic’s context-window documentation explicitly warns that accuracy and recall can degrade as token count grows, a limitation it calls context rot. The same documentation recommends managing long-running agent context rather than treating window size as a cure.

Use the window for a coherent task package. Retrieve the rest by question: “show direct callers of this interface,” “show tests asserting this response,” or “show the current ADR governing this boundary.” For a broader framing of window size versus structure, see Kimi K3 context window vs repo structure.

A Practical Large-Repo Evaluation

Run a private evaluation on a fixed commit and require evidence, not a polished narrative. Pick several real but non-sensitive tasks, grant the same tools and permissions for every run, and preserve the tool trace, plan, diff, tests, and reviewer decision.

Architecture navigation

Give the agent a feature request and ask for the entry points, modules, services, and owners it believes are involved. Review the list against maintainers’ knowledge. Score missing critical surfaces separately from extra suggestions.

Dependency tracing

Ask for direct and indirect consumers of one contract. The answer should cite file paths, symbols, or query results. A claim without a retrieval trace is a lead to inspect, not a dependency finding.

Change planning

Require a plan before edits: intended files, assumptions, migration or compatibility risk, and a stop condition. Anthropic’s Opus 5 prompting guidance recommends a complete task specification for difficult multi-file work; that is compatible with a team-owned plan review, not a replacement for it.

Test discovery

Ask which tests would fail if the intended behavior regressed and why. Then run the selected tests plus the project’s normal gates. A passing narrow test suite does not validate a missing consumer, so reviewers should compare the evidence list with the diff.

Checklist-style editorial diagram for architecture, dependency, plan, and test evidence in a repository evaluation

When External Repository Knowledge Becomes Necessary

External repository knowledge becomes useful when the same relationships must be reused across sessions, tools, or agents: call paths, owners, tested contracts, and design records. A maintained index or graph can retrieve a small, sourced context package instead of repeatedly asking a model to rediscover it.

That layer must expose provenance, revision, and freshness. It should not approve a design, infer permissions, or silently promote generated summaries to authority. Claude Opus knowledge graph explains this boundary; the Graphify hub is the product-level starting point for structured repository context.

FAQ

Can Claude Opus reuse repository knowledge between separate sessions?

Do not assume it can. A context window is request and conversation working memory, and what survives depends on the product surface and workflow. Keep reusable repository facts in versioned project artifacts or a separately governed retrieval layer; reintroduce the relevant revision-bound facts for each new task.

Should teams upload an entire monorepo to a hosted model?

Only after security, data-classification, retention, and vendor-access review. Start with a minimal task package and a non-sensitive pilot repository. Exclude secrets, production credentials, customer exports, and unrelated services by default.

How should proprietary files be excluded from model evaluations?

Write an allowlist before the pilot. Make tool access enforce it, keep the evaluation corpus synthetic or approved, and inspect traces for accidental expansion. A verbal instruction alone is not an access control.

Who owns the final architecture decision after an agent proposal?

The designated engineering owner does. The agent can propose options and cite evidence; reviewers should record the accepted decision and the assumptions that would require reopening it.

When should teams repeat a repository evaluation after model updates?

Repeat it after a model, tool, permission, retrieval, or repository change that could alter behavior. Keep prior runs as dated evidence, not as a permanent capability label.

Conclusion

Claude Opus 5 changes the amount of coherent evidence a coding agent can inspect, not the need for repository evidence. Start with a bounded task, require source-backed architecture and dependency claims, validate the resulting diff with the right tests, and keep the final design decision with accountable reviewers.

Related posts

Latest posts