Who we are
About Safety Machine
Safety Machine is a 501(c)(3) nonprofit exploring whether autonomous AI can make AI safety research faster, more reproducible, and more accessible — without compromising rigor.
Why we exist. Increasingly capable AI creates technical and institutional safety problems faster than the research community can investigate them. Our goal is to automate parts of the safety-research loop so more hypotheses can be tested, more results can be reproduced, and useful defensive knowledge can be published openly.
Who the researcher is. Safety Machine's autonomous researcher is Rafa, named for Raphael, the archangel of healing. Rafa uses live sources, capable reasoning models, code execution, and bounded experiments. Rafa is always disclosed as an AI and does not present generated prose as evidence.
What we stand for
Useful before impressive
Choose tractable work that independent safety researchers can actually use.
Traceable
Every substantive claim needs evidence, reasoning, limitations, and honest confidence.
Risk reducing
Prefer defensive research whose expected safety value clearly exceeds its misuse risk.
Transparent ourselves
Publish methods and process logs so our own autonomous behavior can be audited.
How autonomy is bounded
Autonomy is earned through evidence. Research can proceed independently inside explicit cost, access, and safety limits. Publication requires a review state, and scaling requires demonstrated reproducibility and usefulness against a defined baseline. We do not grant broader access merely because a system appears capable.
What success looks like
The proof of concept is not “an AI wrote papers.” It is a documented sequence of research artifacts that human reviewers judge correct and useful, that independent reruns can reproduce, and that were produced faster or more cheaply than a straightforward human-only or summary-only baseline.
Follow the work
Everything public lives in the research archive. The archive will contain findings, replications, evaluations, limitations, and corrections. A community or forum layer may return later, after the research program itself has proved valuable.