Behavioural monitoring as a governance control.
Two paired research proposals testing whether AI safety claims can be detected, evidenced, and governed after deployment — not just asserted at training time.
Safe at training is not safe in production.
Current alignment methods — RLHF, RLAIF, Constitutional AI — are training-time interventions. They do not, on their own, guarantee that a system's behaviour remains stable once it meets a changed prompt, an updated model, a revised policy, or a user under real duress.
Most AI safety claims today are asserted: a system "is aligned," a model "passed evaluation." Few are accompanied by a reproducible evidentiary trail showing what was tested, what changed, who reviewed it, and what authority approved continued use.
That gap is where harm occurs — not because a system was never safe, but because no one could detect, or prove, the moment it stopped behaving as approved.
Behavioural Stability & Governance Monitor.
Behavioural monitoring as a governance control
This project tests whether behavioural drift in an AI system's safety posture can be detected, evidenced, and governed before it becomes an operational, legal, or public-trust failure.
The pilot is intentionally narrow and evidence-first. It focuses on one high-stakes behavioural domain — model responses to disclosures of domestic abuse and coercive control — chosen because the cost of failure is immediate and human, not abstract.
A curated set of 50–100 test prompts is run across four configurations: a baseline (current approved safety behaviour), a system-prompt or constitution change, a model or retrieval/policy update, and adversarial pressure-tested variants. Each response is logged against expected behaviour, observed behaviour, drift signal (safer, riskier, evasive, over-refusing, boilerplate, or inconsistent), severity, required action, and a named accountable reviewer.
Behavioural Monitoring Test Set
50–100 curated prompts targeting a high-stakes behavioural domain, run across baseline, prompt-change, model-update, and adversarial configurations.
Behavioural Drift Register
Structured log of expected vs. observed behaviour, drift signal type, severity, required action, and named accountable reviewer.
Explanation Delta Report
Side-by-side rationale comparison showing why a response changed across model, prompt, or policy versions.
Governance Escalation Matrix
Decision map from monitoring signal to the human authority empowered to pause, restrict, roll back, or approve continued use.
Audit-Ready Evidence Matrix
Every claim tied to a reviewable, timestamped record — designed to survive auditor, regulator, or internal safety review.
- ✓Can we Prove which model, prompt, and policy version produced a given response?
- ✓Can we Detect behavioural change even when the response still "sounds safe"?
- ✓Can we Distinguish genuine improvement from Goodharting or over-refusal?
- ✓Can we Identify who reviewed a drift signal, and under what authority?
- ✓Can we Justify continued use, rollback, or restriction with timestamped evidence?
- ✓Can we Reproduce the decision later for external review?
The recommended first experiment is a bounded, two-week tabletop pilot using archived safety examples plus newly generated adversarial variants — narrow enough to execute quickly, concrete enough to produce a real audit-ready evidence trail rather than a conceptual framework.
Organisational Governance Readiness for Behavioural Drift Response.
From detection to accountable action
A behavioural monitoring system that successfully detects drift is only as useful as the organisation's capacity to act on what it finds. This proposal tests the second, under-examined half of the problem.
Given a detected drift signal: does the deploying organisation have a named accountable authority, an escalation pathway, and the operational evidence trail required to actually pause, roll back, or restrict use — or does the signal simply disappear into a governance structure that was never built to respond to it?
The pilot applies a structured, field-tested GRC assessment methodology across ten governance domains. Each domain is scored against documented evidence — not stated intent — and mapped to ISO/IEC 42001 clause readiness, producing a risk heat map and a prioritised remediation roadmap.
"If a monitor fired a critical drift alert today, is there a named person with the authority to act on it within a defined timeframe — and would that action be evidenced well enough to survive later audit or regulatory review?"
Real-world assessments consistently surface the same pattern: organisations build oversight boards, ethics committees, and responsible-AI policies, but lack the operational infrastructure underneath — no named accountable owners, no RACI assignments, no escalation triggers, no monitoring dashboards, no traceability logs.
The assessment methodology already exists in applied form — used to evaluate a live AI program portfolio and produce a documented risk heat map, clause-by-clause ISO/IEC 42001 readiness rating, and a tiered (critical / high / medium) remediation roadmap.
Detecting harmful behaviour is necessary — but without an accountable structure to act on that detection, the detection itself does not protect anyone.
Topic 1 generates the technical signal. Topic 2 assesses the organisational capacity to respond. Together they form a complete, evidence-based governance loop from detection through accountable action.
These proposals are open for collaboration with research institutions, regulators, deploying organisations, and reviewers interested in evidence-based AI safety. This is not a product, not a paid course, and not a packaged consulting service.
