GeoVerse Lab
← AI & Computing Sciences Division

AI Safety & Alignment Center

Researches AI alignment, mechanistic interpretability, robustness, and evaluation methods - including RLHF, red-teaming, and scalable oversight - to ensure advanced AI systems remain safe, controllable, and beneficial when deployed in high-stakes scientific and societal settings.

Machine Learning CenterNatural Language Processing CenterComputer Vision CenterHigh-Performance Computing CenterQuantum Computing CenterGenerative AI & Foundation Models CenterReinforcement Learning & Agents CenterAI Safety & Alignment CenterAI for Science CenterRobotics & Embodied AI Center
Research Fields 15
Value Alignment & Cybernetic Control
πŸŽ›οΈ Value Alignment & Cybernetic Control
Value alignment is the collection of training and evaluation techniques, chiefly reinforce…
Preference Learning & Reward Modeling
πŸ‘ Preference Learning & Reward Modeling
Preference learning trains a reward model to predict which of two candidate outputs a huma…
Mechanistic Interpretability of Neural Representations
πŸ” Mechanistic Interpretability of Neural Representations
Mechanistic interpretability reverse-engineers the internal computations of a trained neur…
Adversarial Robustness & Security Evaluation
πŸ›‘οΈ Adversarial Robustness & Security Evaluation
Adversarial robustness is the study of how reliably a trained model maintains correct or s…
Scalable Oversight & Multi-Agent Debate
βš–οΈ Scalable Oversight & Multi-Agent Debate
Scalable oversight studies how a relatively weak evaluator, such as a non-expert human or …
AI Governance & Ethics of Technological Responsibility
πŸ›οΈ AI Governance & Ethics of Technological Responsibility
AI governance is the interdisciplinary study and practice of designing institutions, regul…
Formal Verification & Safety-Critical AI Systems
βœ… Formal Verification & Safety-Critical AI Systems
Formal verification proves, through mathematical deduction rather than testing on a finite…
Existential Risk Analysis & Long-Term AI Safety
⚠️ Existential Risk Analysis & Long-Term AI Safety
Existential risk analysis systematically constructs and evaluates scenarios in which an ev…
Inverse Reinforcement Learning & Mechanism/Incentive Design
🎯 Inverse Reinforcement Learning & Mechanism/Incentive Design
Inverse reinforcement learning infers the reward function that best explains an observed a…
Algorithmic Fairness & Distributive Justice in AI
βš–οΈ Algorithmic Fairness & Distributive Justice in AI
Algorithmic fairness studies how to define, measure and enforce fairness in the decisions …
Trustworthy AI & Human-AI Symbiosis
🀝 Trustworthy AI & Human-AI Symbiosis
Trustworthy AI research designs the interaction protocols, interface elements and organisa…
Safety Engineering & Systemic Risk of Complex AI Systems
🏭 Safety Engineering & Systemic Risk of Complex AI Systems
Systemic risk analysis of complex AI systems examines how failures can emerge from the int…
Machine Ethics & Deontological Constraint Design
πŸ“œ Machine Ethics & Deontological Constraint Design
Machine ethics designs the value systems embedded in an AI system's training and behaviour…
Calibration & Uncertainty Estimation for AI Safety
πŸ“ Calibration & Uncertainty Estimation for AI Safety
Calibration measures whether the confidence a model reports alongside a prediction matches…
Control Theory & Corrigibility of Autonomous Agents
πŸ•ΉοΈ Control Theory & Corrigibility of Autonomous Agents
Corrigibility research designs autonomous agents that reliably accept correction, modifica…