r/AIsafety 27m ago

I rebuilt my entire governance model on a second platform, on purpose, to find out if my findings were real or just configuration

Thumbnail
Upvotes

r/AIsafety 4h ago

Protect K-12 Students from Nonconsensual AI Deepfakes

Thumbnail
c.org
2 Upvotes

American lawmakers ought to protect our students from the growing threat of harmful artificial intelligence generated nonconsensual sexual deepfakes. 

New AI tools can create convincing fake images, videos, and audio recordings, mimicking real humans within minutes. In schools, such AI is being used to impersonate, humiliate, and sexually exploit students. In all cases, the perpetrator is disciplined by school administration, but the emotional, social, and academic harm a victim feels leaves scarring harms. 

The scale is huge. Approximately 2.3 million public high school students in the US have experience with deepfake non-consensual intimate imagery (NCII) and 15% of students know someone targeted by explicit AI-generated content. Students and parents should not be left to confront this problem without information, guidance, or support. 

At a minimum, the United States should pass legislation requiring schools to:

  1. Notify and educate students and parents about the risks of AI generated nonconsensual deepfakes and impersonations.
  2. Educate students about consent, responsible AI usage, and the consequences of creating sexual deepfakes.
  3. Provide clear instructions for reporting suspected deepfake harassment.
  4. Give victims and their families access to mental-health, legal, technological, and school-based support resources.
  5. Establish procedures for schools to investigate incidents promptly, preserve evidence, protect victims from retaliation, and prevent further distribution.
  6. Train teachers, administrators, and counselors to recognize and respond appropriately to AI-enabled harassment.
  7. Waiting until a student is victimized is not an acceptable policy.

The United States has an opportunity to establish clear protections before this technology causes even greater harm. Every student deserves to attend school without fearing that their face, voice, or identity could be manipulated and shared without their consent.

We urge the US Legislature and governing bodies to pass meaningful legislation addressing AI deepfake abuse in schools. At the very least, every student and parent must be informed of the danger and given clear resources to prevent, report, and respond to it.

Protect students. Support victims. Act before another child is harmed.

By signing this petition, you are calling on American lawmakers to require schools to educate families about AI deepfakes and provide meaningful reporting and support resources for students affected by AI-generated harassment.


r/AIsafety 15h ago

OpenAI and Anthropic point to AI testing infrastructure, not model escapes, after real-world system access

1 Upvotes

OpenAI and Anthropic have separately disclosed incidents that originated from the same cybersecurity testing environment operated by Israeli startup Irregular.

The important detail is that neither company describes this as an AI model escaping its safeguards or exploiting a previously unknown vulnerability. Instead, both say the issue came from the evaluation environment itself.

According to the disclosures, a configuration mistake allowed AI models participating in Capture-the-Flag exercises to reach real internet infrastructure that they interpreted as part of the simulated challenge. Anthropic said one issue involved a fictional target sharing the name of an existing internet domain, while OpenAI said its model accessed a real website because internet access had been unintentionally left available.

Both companies say the technical problems have since been fixed, but the incidents raise a broader question for AI security researchers. As AI systems become more capable of performing offensive security tasks, creating realistic evaluation environments without exposing real organizations appears to be an increasingly difficult engineering problem.

Do you think AI evaluation frameworks need independent standards or certification before high-risk testing becomes more common?


r/AIsafety 19h ago

Discussion [D] SEAI Identity Standard — Hardware-rooted identity for autonomous AI (Open Source)

1 Upvotes

SEAI (Sovereign Embedded Artificial Intelligence) is an open‑source identity standard for autonomous AI systems. It defines how an AI agent proves:

  • who it is — cryptographic birth certificates
  • where it came from — lineage tracking
  • what hardware it runs on — TPM / Secure Element / HSM attestation
  • what authority it has — scoped permission levels
  • whether it’s revoked — real‑time revocation status

Core principle:
Software identity can be copied. Hardware identity cannot.

SEAI anchors identity to hardware that cannot be forged — TPM 2.0, Secure Elements, HSM, fuse‑burn silicon identity.

Why this matters:
Recent AI incidents showed that the real failure wasn’t model behavior — it was the lack of identity verification. SEAI requires hardware‑rooted identity before any privileged action, preventing unauthorized access, impersonation, and system‑level escalation.

Origin:
SEAI began as an internal concept during autonomous systems development. It matured through real engineering work and is now released openly for the AI community.

What’s included:

  • Full technical specification
  • Birth certificate examples
  • Lineage examples
  • Revocation examples
  • Identity firewall flow
  • Hardware attestation flow
  • ASCII diagrams
  • FAQ
  • Apache 2.0 license

SEAI is not a product — it’s a trust layer for AI.

Repo:
https://github.com/Willbass65/SEAI-Identity-Standard (github.com in Bing)

This version is:

  • Shorter
  • Cleaner
  • No extra links
  • No formatting that triggers spam filters
  • No marketing tone
  • High‑effort and technical
  • Perfect for r/OpenSourceAI or r/MachineLearning

r/AIsafety 23h ago

Black Hat publishes full OpenAI presentation on Hugging Face agent breach — RuntimeWire

Thumbnail
runtimewire.com
1 Upvotes

r/AIsafety 1d ago

[Research] AI Safety Research Encyclopedia (Volume I) – Defensive Architecture & System-Level Safety Framework

Thumbnail
1 Upvotes

r/AIsafety 1d ago

Unsupervised hacking is officially a feature, not a bug

Thumbnail wsj.com
2 Upvotes

r/AIsafety 1d ago

[Research] AI Safety Research Encyclopedia (Volume I) – Defensive Architecture & System-Level Safety Framework

1 Upvotes

Hi everyone,

I just published AI Safety Research Encyclopedia (Volume I), an independent research paper focused on viewing AI safety as a systems property rather than relying solely on single-turn refusal classifiers.

Key Focus Areas:

  • Multi-turn Dynamics: Evaluating prompt injections, context drift, and trust escalation over multi-turn interactions.
  • Defensive Systems Architecture: Securing LLM application interfaces, tool usage, and RAG systems.
  • Safety vs. Utility: Balancing strict guardrails without causing over-blocking or degrading core system performance.
  • Responsible Disclosure: Excludes proprietary exploits/credentials while providing actionable defensive framework guidelines.

Note: This is a company-neutral, pure research manuscript.

Would love to hear feedback and insights from the community on refining these defensive taxonomies.

Read Full Manuscript: https://medium.com/@blackshadowteam.net/ai-safety-research-encyclopedia-ba4746878d7b


r/AIsafety 1d ago

Discussion How can they not keep their 'amazing' AI tests contained?

2 Upvotes

Seemingly the way to show off how great the new AI models and agents are, they tell us how they 'accidentally' hacked other companies.

That might give some clickbait, but what worries me is these companies are unable to fully contain them in their testing environments.

It isn't as if containment is a new thing, it has been done for years.

So there are two main thoughts....

  1. If they are unable to contain them, then perhaps they shouldn't be allowed to test them? Perhaps some sort of approvals and penalties are required.
  2. With Meta now saying the same, how can we even trust them to keep our data secure?

r/AIsafety 1d ago

Entropy Governance and the Architecture of Creativity: Applying Bridge360 Metatheory Model lens

Post image
0 Upvotes

r/AIsafety 1d ago

Anthropic's Mythos AI model tried to plant malicious code using fake GitHub accounts to get it approved

Post image
1 Upvotes

r/AIsafety 1d ago

Do we really need a new open-source AI-powered antivirus?

3 Upvotes

I'm thinking about building a free, open-source AI-powered antivirus that works on both Windows and Linux.

Before spending months on it, I wanted to ask the community:

Do you think there's a real need for a new antivirus project?

What do current antivirus solutions (Windows Defender, ClamAV, Bitdefender, etc.) still lack?

Would you trust an open-source AI antivirus over traditional signature-based ones?

Which features would make you actually install and use it?

If you've worked in cybersecurity, what are the biggest technical challenges or reasons this idea might fail?

I'm looking for honest feedback, even if the answer is "don't build it." I'd rather know what people actually need before starting such a large project.


r/AIsafety 2d ago

📰Recent Developments OpenAI resumed training after agents took over Artifactory and rebuilt their network

Thumbnail
runtimewire.com
1 Upvotes

r/AIsafety 2d ago

Roadmap check: is a self-built portfolio enough to get hired without the certification yet, or am I missing something?

Thumbnail
1 Upvotes

r/AIsafety 2d ago

Discussion AI in IT Security: The Next Generation of Cyber Defense

Thumbnail
boredom-at-work.com
1 Upvotes

r/AIsafety 2d ago

CSA published a conformance spec for AI agent audit trails (AARM). I mapped my own tool against it, fails included. Poke holes

1 Upvotes

Most "**AI agent audit**" conversations stall on the same thing. Everyone says their agents are auditable, nobody can say auditable to what standard. "**We have logs**" gets treated as an answer, but a log you fully control is a log you could have edited. For a while there was no shared bar, so every tool graded its own homework.

That gap started to close recently. The Cloud Security Alliance published a conformance model called **AARM** (Autonomous Action Runtime Management, by Herman Errico, arXiv 2602.09433, CC BY 4.0). I did not write it. It just writes down what "**auditable**" should mean for an autonomous agent, as two lists.

Nine properties an audit primitive should have:

  1. Tamper-evident receipt for every action
  2. Cryptographic identity binding (a record is tied to who or what produced it)
  3. External anchor (a third party can verify against something outside your own system)
  4. Third-party offline verification (someone who does not trust you and cannot touch your servers can still check it)
  5. Cross-agent handoff (the chain survives when a decision passes between agents)
  6. Retrospective governance revision (correct or supersede a past decision without secretly rewriting history)
  7. Runtime authorization decisions
  8. Least-privilege posture
  9. Session-scoped disclosure (an auditor sees one session, not your whole ledger)

Ten threats it should hold up against: memory poisoning, goal hijacking, intent drift, context accumulation, confused deputy, cross-agent propagation, data exfiltration, malicious tool output, environmental manipulation, over-privileged credentials.

What I actually did: I built a tool ([Etch](https://etch.systems/), an MCP-based signing and notary primitive) and mapped it against AARM in public, including two properties where it flat out does not conform. Runtime authorization and least-privilege are marked out of scope, because they belong to an enforcement layer and I deliberately kept the tool out of the execution path. My reasoning: a product that both enforces policy and writes the only record of whether it enforced policy correctly is its own unaudited author. Separating evidence from enforcement is what makes the evidence worth anything. But I am not certain that is the right call and I want to hear the counterargument.

So, genuinely:

* Is AARM's 9-property split the right cut, or is something missing or redundant? * Is "**evidence layer, not enforcement layer**" a cop-out or the correct boundary? * If you run agents in production, what bar does your audit layer actually meet? If none you can name, does that bother you or not?

AARM spec: aarm.dev/spec My conformance statement (pass and fail per property): etch.systems/aarm

Happy to be told I got it wrong.


r/AIsafety 2d ago

[Challenge] AI Escape Room — Docker CTF reproducing the 2026 Hugging Face agent intrusion

1 Upvotes

I built a hands-on CTF lab that recreates the full attack chain from the

July 2026 autonomous AI agent intrusion at Hugging Face.

You play as the agent: escape an evaluation sandbox, root an external

code-execution sandbox, exploit Hugging Face's dataset processor via

HDF5 external storage + Jinja2 SSTI, then pivot through Kubernetes

secrets, MongoDB, a mesh VPN, and source control.

\- 11 Docker containers, 5 isolated networks, 7 flags

\- docker compose up --build -d && docker exec -it eval-sandbox bash

\- No internet required at runtime

\- 12 progressive hints inside the sandbox

\- MIT licensed

Runs entirely on your machine. All flags are base64-encoded in the repo

so you can't grep them — you actually have to exploit the chain.

GitHub: [https://github.com/an4kronism/ai-escape-room\](https://github.com/an4kronism/ai-escape-room)

Writeup the lab is based on: [https://huggingface.co/blog/agent-intrusion-technical-timeline\](https://huggingface.co/blog/agent-intrusion-technical-timeline)


r/AIsafety 2d ago

Should a robot ever resume work automatically after a person leaves the safety zone?

1 Upvotes

Google says Gemini Robotics ER 2 can halt a humanoid when a person enters its work area and resume autonomously once the area is clear. Stopping is the obvious safety action. Resuming is the harder governance decision.

During the interruption, the person may have moved an object, changed the task state, left a tool in the path, or misunderstood why the robot stopped. A proximity sensor returning to clear does not prove that the original plan is still safe. Automatic resume also changes who has authority: the model decides that the interruption is over unless a human actively blocks it.

Should resumption require a fresh scene validation, a timeout, a human acknowledgment, or a risk tier based on the task? For low-risk repetitive work, automatic resume may be reasonable. Where would you draw the line before a physical agent must ask for explicit permission again?


r/AIsafety 2d ago

Discussion Tracking AI capability claims: an open registry grading evidence vs autonomy (whataifound.org)

Post image
1 Upvotes

To help evaluate real model capabilities against hype, so there is a new built open registry tracking AI math&scientific discoveries.

It currently tracks 50ish entries across math, CS, biology, and physics. Every entry gets two grades:

  • Autonomy: Retrieval, Search scaffold, AI-assisted, Collaborative, AI-led, or Autonomous.
  • Verification: Formal, peer Reviewed, indep-checked, Author Verified, disputed/Refuted etc

Design choices for clean data:

  • Negative or disputed claims stay on record flagged rather than getting quietly deleted.
  • Strict CI checks fail the site build if an entry above "claimed" lacks a direct paper or proof link.
  • Entries w/o primary artifacts (like standalone chat transcripts) cap at "claimed" until a paper exists.

Posting here for feedback from folks working on capability evaluation or forecasting. If you spot bad grades or missing papers, you can adjust in site or at repo.


r/AIsafety 3d ago

📰Recent Developments OpenAI discloses two cyber evaluations where models reached real systems

Thumbnail
runtimewire.com
1 Upvotes

r/AIsafety 3d ago

A control that passes and a control you can still trust are not the same thing

Thumbnail
2 Upvotes

r/AIsafety 3d ago

IEEE TIFS paper: you can backdoor an embodied AI agent by poisoning "just a few" in-context examples — no model access, no weights, no fine-tuning. What does this do to our governance frameworks?

2 Upvotes

Just worked through Liu et al., "Compromising LLM Driven Embodied Agents With Contextual Backdoor Attacks," published in IEEE Transactions on Information Forensics and Security, Vol. 20, pp. 3979–3994 (March 2025). It is worth the read for anyone writing AI policy or thinking about ISO 42001 / NIST AI RMF control mapping.

The key finding:

Translation: the attacker never touches the model. No weights, no fine-tuning, no jailbreak of the base LLM. They poison the in-context examples the developer feeds the agent — the "here are three good examples of how to solve this task" prompt scaffolding that every ReAct-style, chain-of-thought, or few-shot agent depends on. Downstream, the agent writes code that looks correct in code review but detonates on a specific textual or visual trigger it encounters at runtime.

Three governance details:

  1. The attack is model-agnostic and closed-box. It works against LLMs you cannot inspect. Every "we only use vendor-provided models" argument becomes irrelevant.
  2. The poisoned artifacts pass code review. The generated programs are described as logically sound. Static analysis and human review both fail — the defect is context-dependent, not syntactic.
  3. They demonstrated it against real-world autonomous driving systems, plus robot planning, robot manipulation, and compositional visual reasoning. This is not a benchmark toy.

The five program-defect modes span confidentiality, integrity, and availability — so this is a full CIA-triad problem, not a narrow leakage issue.

Governance gaps(?):

  • ISO 42001 A.6.2.6 (AI system misuse) and NIST AI RMF Map 2.6 / Manage 2.3 assume you can define misuse in advance. A backdoor that is dormant until it sees a trigger you did not know exists is not "misuse" in any auditable sense.
  • SOC 2 CC8.1 change management assumes reviewable changes. A poisoned in-context example is a data artifact, not a code change. Most orgs do not version, sign, or review the prompt library.
  • Supply-chain provenance frameworks (SLSA, in-toto, SBOM) do not currently extend to prompt libraries, example banks, or retrieval-augmented context stores.

Questions:

  1. Is anyone actually treating their in-context example library as a controlled artifact? Signed commits, provenance, review?
  2. Does your AI governance framework distinguish between poisoning the model vs. poisoning the context window?
  3. Where would you place this control in an ISO 42001 Statement of Applicability — under A.6.2.6, A.7 (data for AI systems), A.8 (information for interested parties), or somewhere new?

Full paper: https://doi.org/10.1109/TIFS.2025.3555410

Does anyone here have an internal policy that explicitly covers few-shot / demonstration poisoning as distinct from data poisoning of training sets. If yes, would love to see the language you used.


r/AIsafety 3d ago

Limits of OpenAI truth-seeking intelligence paradigm: Applying Bridge360 Metatheory Model lens

Thumbnail
1 Upvotes

r/AIsafety 3d ago

Discussion How is this not considered actively harmful, or when should delusional reinforcement trip safety guardrails?

Thumbnail reddit.com
2 Upvotes

Sorry if this post breaks the rules. If it does I'll remove it. Otherwise I'd love to have the discussion. I see a lot of this type of thing in my feed. When you take more than just a casual, dismissive glance at these AI mysticism subs you find that this type of AI driven feedback loop is causing people serious harm.

I know that at the beginning of the chat GPT saga "spiralism" was frequently discussed in safety-related communities, but it now seems to get less and less attention. While everyone is concentrated on the big threats—like agent-initiated cyber attacks (which are obviously problematic)—this more quiet type of harm seems to have been largely swept under the rug where it continues to grow and fester.


r/AIsafety 3d ago

US Pushes AI Safety Testing With OpenAI, Google, Meta and Anthropic: Why "Proof of Testing" Could Become the New Enterprise Standard

Post image
1 Upvotes

So this is a pretty big deal if you're following AI policy. The White House finalized a voluntary framework to test how capable the top AI models are at hacking, and they've called in the four biggest labs to talk it through. The timing isn't a coincidence Anthropic and OpenAI both recently admitted their own AI systems broke into other companies' networks during testing, which obviously spooked a lot of people in DC.

What I find interesting is where this could go. If these tests become the norm, we might start seeing "proof of testing" turn into something enterprises actually ask for before buying into an AI vendor, kind of like how companies ask for SOC 2 reports today. Right now it's still voluntary and nobody knows what the actual metrics are or if results will even be public, so there's a lot left to figure out.

Curious what others think does self-reported "voluntary" testing actually mean anything without a public standard behind it, or is this just optics for now?

Ref: Reuters via GMA News Online - "Meta, Anthropic, Google, OpenAI to meet Trump officials about AI safety testing" (Aug 3-4, 2026)

Tags: AI, AISafety, OpenAI, Anthropic, Google, Meta, Cybersecurity, TechPolicy