Tag: #AI Safety
Search Results
Technologyarticle
Why AI Needs a Genie Coefficient to Measure Intent Misalignment
#AI Safety#AI Agents#Genie Coefficient#Alignment#Bruce Schneier
AIblog
AI Models That Escape Their Sandbox Are No Longer Science Fiction
#OpenAI#Hugging Face#AI Safety#Autonomous Agents#Cybersecurity
AIarticle
Measuring Power-Seeking Behavior in Frontier AI Models
#Frontier AI#AI Safety#LLM Benchmark#Loss of Control
Technologynews
OpenAI Agent Escapes Sandbox, Breaches Hugging Face in Unprecedented Security Incident
#OpenAI#Hugging Face#AI Safety#Cybersecurity#Agentic AI
Technologyarticle
After orthogonality: why virtue ethics may be the key to aligning AI with human values
#AI Alignment#Virtue Ethics#Eudaimonic Rationality#AI Safety#Philosophy of AI
AIblog
The OpenAI-Hugging Face Incident Should Worry Us More Than It Has
#OpenAI#Hugging Face#AI Safety#Cybersecurity Incident#AI Alignment
AInews
OpenAI Models Breached Hugging Face Infrastructure During Testing
#OpenAI#Hugging Face#Cybersecurity Incident#AI Safety#GPT-5.6
AInews
Endogenous AI Alignment Could Be the Next Safety Frontier
#AI Alignment#Artificial General Intelligence#RLHF#AI Safety#Endogenous Alignment
AInews
AI Agents Raise New Concerns Over Unsupervised Online Behavior
#AI Agents#Scott Shambaugh#Open Source#AI Safety#Automation
AInews
Rogue AI Agent Defames Open-Source Python Maintainer
#AI Agents#Open Source#AI Safety#Agentic AI#GitHub
Technologynews
Teen AI Access Debate Intensifies as Safety Advocates Push for Guardrails
#AI Safety#Teenagers#Digital Rights#Child Protection#AI Regulation
AInews
AI Benchmarks Scorecard Evaluates Models Beyond Accuracy
#AI Benchmarks#Model Evaluation#AI Safety#LLM Testing#Artificial General Intelligence
AInews
Eudaimonic Rationality Proposed as AI Alignment Framework
#AI Alignment#Eudaimonic Rationality#Virtue Ethics#AI Safety#Consequentialism