AI safety is the practice of designing, building, testing, and running AI systems so they do what they are supposed to do and do not cause unintended harm.
It covers bugs in a self-driving perception stack and also how a language model might sway public opinion. Practitioners mix computer science, engineering, psychology, and ethics to write verification methods and controls. Models now sit in medical diagnosis, trading, and recommendation. Mistakes include misdiagnoses, market crashes, and harmful misinformation.
One error can cascade through linked services and hit millions of users.
Governments, companies, and civil groups fund standards, audits, and certification. Researchers work on formal verification, incentive design, and interpretability so behavior stays within bounds as capability grows. Putting those checks in place now is cheaper than repairing damage after a failure in production.
Safety work is concrete: test a perception stack so a car does not miss a pedestrian, and test a language model so it does not mass-produce harmful falsehoods. Methods include formal verification, incentive design, and interpretability. Deployment in clinics, markets, and feeds means a bug can cascade.
Audits, standards, and certification programs are how governments and companies try to catch that before millions of users are hit. Building the tests while models are still bounded is cheaper than a recall after a crash, a flash crash, or a misinformation wave. The 2016 "Concrete Problems" paper listed reward hacking, safe exploration, and distributional shift. Those problems are still open.
AI Safety Interactive Lab
Adjust safety measures for different AI systems and observe how they affect risk levels. See how thorough safety practices reduce the likelihood of harmful incidents.
Select AI System
Safety Measures
Risk Assessment
Potential Risks for Self-Driving Car
Safety Tip: AI safety requires a multi-layered approach. No single measure is sufficient - thorough safety comes from combining testing, monitoring, controls, and verification.