FuguReport

Summary

This week's AI safety research emphasizes the shift from broad concern about AI harms toward structured governance and quantitative risk-modeling frameworks. Representative papers highlight that risks such as hallucinations, AI-assisted cyber offense, deepfakes, and psychological harms from conversational agents are already manifesting, and that governance must link technical capabilities to measurable real-world harm.

Situation

As AI systems—especially large language and conversational models—spread across education, healthcare, law, public services, and personal support, the representative papers frame safety problems as immediate rather than speculative. They document risks including prompt-injection attacks, deceptive synthetic media, hallucinations in high-stakes settings, AI-enabled cyberattacks that raise attacker efficiency and reach, and psychological harms linked to anthropomorphic conversational agents.

A common baseline across these works is that current safety practice remains fragmented: technical research often isolates single failure modes, legal and ethical discussions stay high-level, and benchmark-based capability thresholds do not directly measure harm. The emerging direction is a more integrated governance approach across the AI lifecycle—combining intrinsic security, misuse and privacy concerns, and broader social and psychological impacts—supported by quantitative risk modeling and context-sensitive taxonomies.

Infographic (English)

AI Governance and Safety situation infographic

Progress

Prioritization of Risks from Artificial Intelligence: A Delphi Study of 272 International Experts <See Details on Fugu-MT>

A three-round Delphi study of 272 international experts quantifies priorities across 24 AI risks, sector vulnerabilities, and actor responsibilities. This provides a structured, empirical severity ranking rather than relying on broad qualitative risk discussions.

Misaligned AI as a New Insider Risk <See Details on Fugu-MT>

This paper reframes misaligned AI in high-stakes deployments as an insider-risk problem and recommends continuous evaluation and monitoring. It extends established insider-risk policy tools to deployed AI models, addressing a gap in current governance practice.

Risk Assessment of Autonomous Driving: Integrating Technical Failures, Ethical Dilemmas, and Policy Frameworks <See Details on Fugu-MT>

This paper applies integrated AI risk assessment to autonomous driving by linking technical failures, ethical dilemmas, and policy frameworks. It grounds the cross-lifecycle governance approach in a concrete safety-critical domain rather than discussing integration only in general terms.

AICompanionBench: Benchmarking LLMs-as-Judges for AI Companion Safety <See Details on Fugu-MT>

AICompanionBench releases real-world Replika conversations annotated across nine detailed harm categories for evaluating companion-AI safety. It converts psychological and relational safety concerns into a public dataset and evaluation benchmark instead of leaving them as conceptual risks.

A Framework for Measuring Appropriate Reliance on Set-Valued AI Advice <See Details on Fugu-MT>

This work introduces a formal framework for measuring appropriate reliance on AI advice in sequential judgment and oversight settings. It captures nuanced human-AI oversight behavior that standard reliance metrics miss, supporting more quantitative safety evaluation.

Outlook

Outlook Summary

AI safety research is shifting from broad governance language toward operational risk management that is measurable, continuously updated, and tied to concrete harms. This week’s work on risk prioritization, deployed-model monitoring, and safety-critical assessment supports a move toward quantitative risk models, continuous evaluation, and lifecycle governance. A second direction is stronger measurement of subtle human-facing failures, not just classic benchmark scores. New work on companion-AI harms and reliance metrics points to public evaluation settings for psychological risk, oversight behavior, and context-sensitive harms. This matches a broader need for adaptive real-world benchmarks, uncertainty-aware evaluation, and safer refusal or self-correction when systems face risky inputs.

Infographic (English)

AI Governance and Safety outlook infographic

Three-Year Movement

The standard path extends the current shift toward operational AI risk management through a HACCP-style mechanism. HACCP is a food-safety method that identifies hazards, sets control limits, monitors key points, and records corrective action. The AI version is a harm-linked assurance ledger: a living record that connects specific hazards to evidence, thresholds, owners, and remediation steps.

In the first year, the main work is making ledger entries credible. Research should turn cyber misuse, companion-AI harms, and human reliance into measures that can support real release decisions. A key monitoring cue is whether release gates or procurement templates begin requiring evidence across several control points, rather than accepting broad benchmark summaries. In the second year, if that cue appears, the agenda moves toward common schemas for assurance ledgers. Researchers would define evidence grades, incident categories, and uncertainty labels that can be compared across systems and settings. Review groups would ask not only how a model performed before launch, but what is monitored after launch and who must act when a control fails.

By the third year, this becomes a more mature but contested assurance layer. Governance platforms would connect evaluation results, monitoring, and incident intake so that approvals can depend on updated harm evidence. Higher-risk deployments would face renewal decisions based on what has happened since release, not only on static claims made earlier. The feedback loop is the mechanism: better monitoring improves incident categories, better categories improve benchmarks, and better benchmarks make assurance records more useful. The caveat is that AI behavior can change across contexts faster than many physical production processes. A disconfirming cue would be evidence that control-point metrics mainly create a sense of order without predicting lower harm.

The contender path accepts the move toward operational AI safety, but adds a workforce bottleneck. The key mechanism is scarce expert capacity. Strong evaluation depends on people such as red-teamers, cyber-risk modelers, and psychological-harm assessors who can apply methods carefully. If deployment grows faster than this workforce, the main problem becomes not only what to measure, but how often credible measurement can be done.

In the first year, research would start measuring the bottleneck itself. It would ask how many qualified evaluators exist, how long reviews take, and which tasks can be automated without losing context. Work on companion-safety evaluation, cyber-misuse modeling, and reliance metrics would continue, but the practical question would be how much expert time is needed for credible deployment gates. A monitoring cue is a clear mismatch between required evaluation work and available labor, such as review queues stretching for months or flaw reports missing response targets.

By the second year, persistent scarcity would push governance toward tiers. The analogy is distributed-energy interconnection, where smaller systems often use pre-certified components instead of full custom review. In AI, that means low-risk systems may use lighter checks, while higher-risk deployments receive deeper independent assessment. Platform providers could become governance intermediaries by offering pre-evaluated endpoints, safety documentation, and monitoring dashboards.

By the third year, the weak point is context. A system may satisfy a standard configuration yet cause harm when used in a setting that the review did not cover. Research would then move toward adaptive certification that updates when monitoring signals or threat conditions change. The caveat is that this path weakens if automated evaluation becomes much cheaper and more reliable, or if early failures discredit pre-certified configurations before they become common.

The maybe path turns safety evidence into an external assurance system. Its mechanism is pressure from buyers and assurance bodies that want reusable proof of responsible deployment. The analogy is the old boiler inspection regime, where inspection certificates and engineering codes reinforced safer practice. In AI, the comparable artifacts are evaluation reports, telemetry records, and structured flaw reports.

In the first year, research would test whether these artifacts are strong enough outside the lab. Companion-harm labels would be checked across user groups, cyber-uplift estimates would be tested against weak assumptions, and risk records would be compared across independent teams. If procurement language starts naming specific evaluations, the field would also study whether the tests themselves can be gamed. A monitoring cue is concrete friction, such as a buyer requiring a prioritized risk record before adoption.

In the second year, the scenario advances only if the forcing loop starts to work. Required telemetry would improve incident data, better data would support more meaningful attestations, and recognized attestations would become more valuable in procurement. Research would shift from single benchmarks to conformity assessment, meaning agreed methods for deciding whether a system meets a safety requirement. Because AI systems can change after release, these methods would need to combine static tests with live monitoring.

By the third year, continuous attestation becomes the central idea. Instead of asking whether a model passed one test on one day, organizations would ask whether the deployment has ongoing telemetry, update-triggered re-evaluation, and clear responsibility for fixes. Higher-risk uses in public services, healthcare, and education would be more likely to rely on managed deployment packages with recognized attestations. The caveat is that the path can fail if rules stay only principle-based, benchmarks are visibly gamed, or evaluator capacity remains too thin. In that case, assurance artifacts would remain pilots rather than becoming a broad deployment norm.

1-Year / 3-Year Research-Application Infographic

Mixed-scenario 1-year/3-year research/application infographic

References

This page was created using generative AI such as GPT-5, Claude Opus 4, Gemini 3, Grok 4, Fugu Ultra, Gemini 3.1 Flash Image, GPT Image 2, and their higher-end successor versions. No guarantee can be made regarding its contents.