AI-written code in regulated industries: what an audit trail needs

View as Markdown
Ask AI
Share

Singapore’s MAS Technology Risk Management Guidelines expect financial institutions to control change management, access and third-party risk. AI agents that write production code touch all three. If an agent writes a change that later causes an incident, the institution needs to be able to explain how that change was produced, who approved it and on what evidence.

Evidence, not assurances

We think every agent-produced change should come with its own evidence: the session log, the diff, the tests it ran and the person who approved it. That is the standard we are designing Stratamend around, so that risk teams can trace how any line of code came to exist.

What we plan to record

For every agent session, the audit record should answer five questions.

  • What was the agent asked to do? The task, the module in scope, and the plan it was given, including the version of any instructions and policies in force.
  • What could it do? The tools it was allowed to use, the directories it could edit and the commands it could run, captured as configuration rather than described after the fact.
  • What did it actually do? Every tool call: files read and edited, commands run and their exit codes. Attempts that a policy hook blocked are recorded too, because they show the boundaries working.
  • What evidence supports the result? The golden-master and unit test results for the final change, plus static analysis output.
  • Who approved it? The reviewer, the time of approval and the pull request in which the change was merged.

Change management

Agent changes should flow through the same change process as human changes, not around it. In our design an agent can only open a pull request. Branch protection, required reviews and CI checks stay exactly as the institution has configured them. The audit record is attached to the pull request, so the reviewer sees the evidence in the place they already work.

Access control

An agent is a privileged actor and should be treated like one. That means least-privilege credentials scoped to a single repository and task, no access to production data, and secrets kept out of the agent’s environment entirely. Hooks that block reads of credential files and reject diffs that look like secrets add a second layer.

Third-party risk

Using an AI model is a third-party dependency. Institutions will want to know where source code is processed and under which agreements. This is one reason we are designing Stratamend to run inside the customer’s own cloud account and call Claude through Amazon Bedrock or Google Cloud Vertex AI in a chosen region, so that the model provider relationship sits within cloud agreements the institution already manages.

Retention and readability

An audit trail only helps if someone can read it. Raw agent transcripts are long. We plan to keep the full record for retention purposes and generate a short, structured summary for reviewers and auditors, with every summary line linking back to the underlying entries.

None of this is exotic. It is the discipline regulated institutions already apply to human developers, made explicit for a new kind of contributor.