AI Safety is a field of research focused on ensuring that artificial intelligence systems behave predictably, avoid causing harm, and remain aligned with human interests. It spans technical alignment, risk mitigation, and the study of existential risk from advanced systems.
Defines the safety alignment and security constraints of user-facing systems during policy creation, alignment training, and jailbreak defense; implementing AI Safety helps builders isolate instructions from injection exploits.
AI safety is a technical research field focused on preventing artificial intelligence systems from behaving in ways that harm humans or act counter to designer intent. It spans short-term risks, such as model hallucinations, toxic outputs, and jailbreaks, as well as long-term risks like loss of control over highly capable autonomous agents. AI safety research encompasses alignment techniques (like RLHF and Constitutional AI), rigorous evaluation benchmarking, and the development of robust guardrails.
The difficulty of ensuring that an AI system's goals match human values, especially when the model is smarter than its creators.
Restrictions placed on AI access to critical infrastructure, networks, or dangerous information to prevent misuse.
Reference this definition in your articles, research, or documentation to credit this source:
OpenAI's new Astra model will use "recurrent depth," a technique that allows the model to operate outside of the sequential thinking that characterizes most...
Trump's secret reviews of frontier AI model may hide corruption, lawsuit says.
OpenAI is calling for California to strengthen SB 53, an AI safety bill that the company previously opposed.
OpenAI reaffirms Zero Data Retention for eligible API customers and previews Private Safety Processing for advanced AI safety without compromising data privacy.
OpenAI reaffirms Zero Data Retention for eligible API customers and previews Private Safety Processing for advanced AI safety without compromising data privacy.
At Ai4, three of the world's most respected AI experts - Geoffrey Hinton, Fei-Fei Li, and Andrew Ng - debated regulation, open source access, and how America...
Weng previously served as the VP of AI Safety Research at OpenAI.
Open source software is a critical pillar of the global economy. It underpins cloud computing, financial services, manufacturing, telecommunications...
Our approach to AI policy and political advocacy, transparency, support for thoughtful regulation and AI safety, and that no outside political group speaks on...