A language model-based agent can adapt to a new interpretation of credit rules, but doing so without boundaries risks knocking out the very control mechanisms required in a regulated sector. If the agent rewrites itself, it becomes difficult to determine what changed, why, against what data it was verified, and who authorized its rollout to production.

This is the problem addressed by a new paper published on arXiv by Ravil Akhtyamov, presented on October 7, 2026. The proposal does not aim to make the language model autonomously modifiable: it freezes the model weights and allows the agent to evolve solely through the so-called execution harness. In other words, operational instructions, tool-invocation logic, and the combination of elementary functions can change; the underlying model remains untouched.

The distinction is particularly relevant in credit assessment workflows, where an automated or assisted decision can impact access to financing and must be auditable after the fact. The paper's core idea is to turn every adaptation into a verifiable artifact: an identifiable modification linked to the cause that motivated it, accompanied by a test, and subject to a formal approval process.

Modifying behavior without making the system opaque

In the debate around AI agents, the ability to self-improve is often described as an operational advantage: the system observes recent results, identifies a problem, and adjusts its strategy. In a credit workflow, however, that same dynamic can produce consequences that are difficult to manage. A fix that reduces errors within a limited sample could, for instance, achieve that outcome simply by loosening checks and letting more problematic cases pass through.

Akhtyamov therefore proposes a technical and organizational perimeter. The agent can suggest changes to the wrapper that directs it: the instruction text, the conditions of the tool-calling logic, and the order and composition of predefined primitives. It cannot directly intervene in the model weights. This choice makes it possible to describe a change as a diff—that is, a specific difference relative to the previous configuration—instead of treating it as an internal transformation of the model that is not easily inspectable.

The process is organized into two loops and includes a single admission gate prior to deployment. The gate logs the operation's elements in a hash chain: a method useful for linking records so that subsequent alterations are detectable. In the framework outlined by the author, traceability is not a documentation task added after the fact, but a necessary condition for the change to be accepted.

This architecture also separates two planes that tend to overlap in agentic applications: the ability to propose a modification and the authority to deploy it. An agent can explore alternatives, but it should not push them into a credit process simply because they appear better on the most recent signals. Verification against explicit criteria is required, including the effects on cases the system should have flagged.

The results: fewer false positives is not enough

The paper evaluates the mechanism in simulation, not on an actual credit platform. This is an essential caveat when interpreting the results: neither the simulated agent nor the component proposing the changes are production LLMs. For the latter, a seeded search system was used, designed to measure the procedure's behavior under controlled conditions.

The trial considers three families of supervision reinterpretation, each at three intensity levels, with ten runs per configuration. In this scenario, the gate admitted 144 changes out of 7,449 candidates. According to the study, none of the accepted changes degraded the error rate on the historical data kept separate for validation. In low- and medium-intensity cases, the system brought the false positive rate back to the level of an ideal benchmark without increasing missed alerts.

The most critical comparison, however, concerns the admission criteria. Replacing the gate with more permissive logic that focuses on reducing errors in recent traces led the same cycles to accept 309 harmful modifications. In 49 out of 90 runs, missed alerts remained above 10%. False positives dropped, but partly because the filter had been relaxed: an apparently better outcome that concealed a deterioration in oversight.

In credit evaluation, this distinction has practical implications. Optimising an isolated metric, especially when measured on recent data, does not equate to improving the overall quality of a control. A system can produce fewer false alerts while simultaneously missing more cases requiring review. Evaluation must therefore incorporate a history relabelled under the new rule interpretation, rather than relying solely on labels established before the change.

When adaptation requires changing tools

The study also distinguishes between different types of rule shifts. Parametric and scope shifts were corrected locally by adjusting the harness's permitted configurations. By contrast, a structural change required replacing a primitive. This step demonstrates that constraints are not merely limitations, but also a way of making it clear when fine-tuning is no longer enough.

However, the proposed solution does not solve every case. At maximum structural change intensity, the gate's fixed tolerance blocked the correct replacement of the primitive in half of the runs across different seeds. This is a well-known trade-off in monitored systems: overly cautious thresholds can prevent necessary fixes, while overly lenient thresholds allow changes that improve one metric at the expense of others.

Another finding is just as significant. When candidates were evaluated against labels from before the interpretation change, the gate rejected all of them. To truly adapt to a new rule, then, the system must incorporate that new interpretation into the evaluation of historical data. It is not enough to ask the model to perform more efficiently against a past shaped by different criteria.

The regulatory question remains open

The author links the described mechanisms to the requirements of the European AI Act applicable to high-risk credit scoring systems: risk management, event logging, documentation, human oversight, and accuracy and robustness requirements serve as implicit benchmarks for a design that preserves evidence of every change. The paper also notes that the US guidelines on model risk published in April 2026 exclude agentic AI from their scope, leaving a less defined regulatory area.

For now, this does not demonstrate that a self-adaptive agent is ready to make decisions on real-world credit. The contribution is an experimental framework, circumscribed and based on simulations. Testing on operational data, comparison with language models actually employed as proposers, handling of incomplete or biased data, and scrutiny by supervisory authorities are still missing. The author has made the code itself, configurations, and single-run outputs available—a useful element for reproducing and challenging the results.

The direction indicated, however, is clear: in regulated environments, automatic adaptation cannot be regarded as an intrinsic property of the model. It must be treated as a controlled modification to the software that puts it to work. For banks, fintechs, and AI tool providers, the challenge will not merely be choosing an agent capable of updating a decision-making flow, but demonstrating that every update remains understandable, testable, and revocable.

Sources