Articles Tagged “AI Safety”
10 articles found

ChatGPT for Teens Adds Study Mode and Quiet Hours
OpenAI's ChatGPT for Teens launched for ages 13-17 with Study Mode, scheduled Study Hours, parental Quiet Hours, and stronger content limits.

Apple Music Will Show Made With AI Labels This Year
Apple Music will make its AI Transparency Tags mandatory and surface Made With AI labels to listeners later in 2026, shifting from removal to disclosure.

OpenAI Model Security Adds Sandboxes and 30-Min Alerts
OpenAI published a new internal security bar: hard sandboxing, activation classifiers on sampled tokens, and a 30-minute alert-or-pause rule.

GLM-5.3 Posts an 84.5% CyberGym Cyber Defense Score
Z.ai's GLM-5.3 lifts CyberGym from 77.2% to 84.5% on post-training alone, and the team is holding weights back two weeks for safety hardening.

Fable 5 Biology Safeguards Cut False Positives by 85%
Anthropic retuned Fable 5's biology classifier, cutting biology fallbacks roughly 85% and total fallbacks 67% on Claude.ai while keeping dual-use limits.

Shieldstral Runs Multimodal Safety on One 16GB GPU
Mistral's Shieldstral is a 3B open-weight safety classifier covering 12 languages and images, taking plain-language policies at inference on a 16GB GPU.

Google DeepMind's AI Control Roadmap Charts a Safer Path for AI Agents
Google DeepMind published its AI Control Roadmap on June 18, 2026 — a defense-in-depth framework for safely deploying AI agents that keeps systems secure even when alignment is imperfect.

Claude Fable 5: Anthropic Brings Mythos-Class AI Power to Everyone, Safely
Anthropic released Claude Fable 5 on June 9, 2026 — its most powerful generally available model yet, a Mythos-class AI with built-in safety fallbacks.

MIT Researchers Develop a New Metric That Catches Overconfident AI Models Before They Hallucinate
A new uncertainty measurement technique from MIT combines self-consistency checks with cross-model disagreement to flag when LLMs generate confident but incorrect responses.

NVIDIA Open-Sources NemoClaw — A Security-First Stack for Deploying Autonomous AI Agents on Any Hardware
Built on the OpenClaw platform, NemoClaw bundles Nemotron models with sandboxed execution and privacy controls, enabling secure AI agent deployment from RTX laptops to DGX clusters.
