GeoVerse Lab
โ† AI & Computing Sciences Division

Reinforcement Learning & Agents Center

Develops reinforcement learning theory, autonomous agent architectures, and sequential decision-making systems - spanning deep RL, model-based world models, multi-agent systems, and LLM-based agents - for scientific exploration, experiment design, and adaptive control.

โš™๏ธ GeoVerse System Administration Division๐Ÿ–ฅ๏ธ AI & Computing Sciences Division๐Ÿ”ฌ Basic Sciences Divisionโšก Intelligent Geophysical Exploration Division๐Ÿ›ข๏ธ Resource & Energy Engineering Division๐ŸŒ Applied Geoscience Solutions Division๐Ÿ’ผ Economics, Policy & Strategy Division๐ŸŽ“ Education & Training Development Division๐Ÿš€ Innovation & International Collaboration Division
Machine Learning CenterNatural Language Processing CenterComputer Vision CenterHigh-Performance Computing CenterQuantum Computing CenterGenerative AI & Foundation Models CenterReinforcement Learning & Agents CenterAI Safety & Alignment CenterAI for Science CenterRobotics & Embodied AI Center
Richard Bellman๐Ÿ”‘
โญ Center Chief
Richard Bellman
Center Head
๐Ÿ’ก Dynamic programming, Bellman equation underlies RL
๐Ÿง‘โ€๐Ÿ”ฌ
๐Ÿ”‘
1941โ€“1997
Harry Klopf
Researcher
๐Ÿ’ก Proposed the 'hedonistic neuron' hypothesis โ€” that individual neurons act to maximize local reward-like signals โ€” directly inspiring Sutton and Barto's actor-critic architecture, which separates a policy ('actor') from a value-estimating critic
๐Ÿ–ฅ๏ธ AI & Computing Sciences DivisionReinforcement Learning & Agents Center
๐Ÿง‘โ€๐Ÿ”ฌ
๐Ÿ”‘
1919โ€“1997
Yakov Tsypkin
Researcher
๐Ÿ’ก Founded the theory of adaptive and learning control systems, showing how a controller can adjust its own parameters online based on observed system behavior โ€” directly foundational to adaptive control-theoretic reinforcement learning agents that tune their policies in response to a changing environment
๐Ÿ–ฅ๏ธ AI & Computing Sciences DivisionReinforcement Learning & Agents Center
๐Ÿง‘โ€๐Ÿ”ฌ
๐Ÿ”‘
1924โ€“1976
Daniel Berlyne
Researcher
๐Ÿ’ก Founded the experimental psychology of curiosity, showing that novelty, uncertainty, and complexity themselves drive exploratory behavior independent of external reward โ€” the foundational psychological theory that intrinsic-motivation and curiosity-driven exploration bonuses in RL directly formalize
๐Ÿ–ฅ๏ธ AI & Computing Sciences DivisionReinforcement Learning & Agents Center
Gustav Elfving๐Ÿ”‘
1908โ€“1984
Gustav Elfving
Researcher
๐Ÿ’ก Founded the theory of optimal experimental design, determining how to choose experimental conditions (the 'Elfving set') to maximize the statistical information gained per experiment โ€” foundational to Bayesian optimal experimental design for autonomous scientific exploration
๐Ÿ–ฅ๏ธ AI & Computing Sciences DivisionReinforcement Learning & Agents Center
๐Ÿง‘โ€๐Ÿ”ฌ
๐Ÿ”‘
1927โ€“2021
Peter Whittle
Researcher
๐Ÿ’ก Developed the theory of restless bandits and Whittle indices, generalizing multi-armed bandit theory to more realistic settings where unobserved arms continue to evolve โ€” foundational to modern exploration-exploitation algorithms balancing information gathering against exploitation of known-good options
๐Ÿ–ฅ๏ธ AI & Computing Sciences DivisionReinforcement Learning & Agents Center
๐Ÿง‘โ€๐Ÿ”ฌ
๐Ÿ”‘
1925โ€“2014
Harold W. Kuhn
Researcher
๐Ÿ’ก Developed the theory of extensive-form games (game trees with sequential moves and information sets), providing the exact formal representation of sequential, multi-step decision-making under uncertainty that both game-tree search and sequential multi-agent RL directly operate on
๐Ÿ–ฅ๏ธ AI & Computing Sciences DivisionReinforcement Learning & Agents Center
๐Ÿง‘โ€๐Ÿ”ฌ
๐Ÿ”‘
1927โ€“1992
Allen Newell
Researcher
๐Ÿ’ก Co-developed the General Problem Solver and means-ends analysis, the foundational method of decomposing a complex goal into a hierarchy of subgoals โ€” directly foundational to hierarchical reinforcement learning's options and subgoal frameworks
๐Ÿ–ฅ๏ธ AI & Computing Sciences DivisionReinforcement Learning & Agents Center
Albert Bandura๐Ÿ”‘
1925โ€“2021
Albert Bandura
Researcher
๐Ÿ’ก Founded Social Learning Theory, demonstrating that organisms learn complex behaviors by observing and imitating others (the famous Bobo doll experiments) without direct trial-and-error reward โ€” the founding psychological phenomenon that imitation learning and behavior cloning in RL directly formalize
๐Ÿ–ฅ๏ธ AI & Computing Sciences DivisionReinforcement Learning & Agents Center
Arturo Rosenblueth๐Ÿ”‘
1900โ€“1970
Arturo Rosenblueth
Researcher
๐Ÿ’ก Co-founded cybernetics with Wiener and Bigelow, formalizing the theory of goal-directed, feedback-driven behavior in machines โ€” the founding theoretical framework for agents that pursue goals via perception-action-feedback loops, directly relevant to LLM-based autonomous agents using tools and feedback
๐Ÿ–ฅ๏ธ AI & Computing Sciences DivisionReinforcement Learning & Agents Center
๐Ÿง‘โ€๐Ÿ”ฌ
๐Ÿ”‘
1901โ€“1990
Arthur Samuel
Researcher
๐Ÿ’ก Built the first self-improving game-playing program (checkers) that learned by playing against itself and evaluating board positions via minimax search, coining the term 'machine learning' and founding the self-play tree-search paradigm that MCTS and AlphaZero-class systems directly descend from
๐Ÿ–ฅ๏ธ AI & Computing Sciences DivisionReinforcement Learning & Agents Center
David Blackwell๐Ÿ”‘
1919โ€“2010
David Blackwell
Researcher
๐Ÿ’ก Proved foundational optimality and convergence results for Markov decision processes (Blackwell optimality), establishing the rigorous mathematical conditions under which dynamic-programming-based sequential decision policies are provably optimal โ€” extending and rigorizing Bellman's framework
๐Ÿ–ฅ๏ธ AI & Computing Sciences DivisionReinforcement Learning & Agents Center
๐Ÿง‘โ€๐Ÿ”ฌ
๐Ÿ”‘
1918โ€“2016
Jay Wright Forrester
Researcher
๐Ÿ’ก Founded System Dynamics, the discipline of building explicit feedback-loop simulation models of complex systems to predict and reason about their behavior โ€” the direct conceptual and terminological ancestor of the 'world models' that model-based reinforcement learning agents learn and plan within
๐Ÿ–ฅ๏ธ AI & Computing Sciences DivisionReinforcement Learning & Agents Center
Lloyd Shapley๐Ÿ”‘
1923โ€“2016
Lloyd Shapley
Researcher
๐Ÿ’ก Founded the theory of stochastic games, generalizing Markov decision processes to multiple interacting agents with competing or cooperative objectives โ€” the exact mathematical framework underlying multi-agent reinforcement learning
๐Ÿ–ฅ๏ธ AI & Computing Sciences DivisionReinforcement Learning & Agents Center
Magnus Hestenes๐Ÿ”‘
1906โ€“1991
Magnus Hestenes
Researcher
๐Ÿ’ก Founded modern numerical methods for the calculus of variations and constrained optimization (conjugate gradient method, augmented Lagrangian methods), providing the gradient-based optimization machinery that policy gradient reinforcement learning algorithms directly apply to optimize policy parameters
๐Ÿ–ฅ๏ธ AI & Computing Sciences DivisionReinforcement Learning & Agents Center
Ivan Pavlov๐Ÿ”‘
1849โ€“1936
Ivan Pavlov
Researcher
๐Ÿ’ก Discovered classical conditioning โ€” that an organism learns to associate a predictive stimulus with a delayed reward through repeated temporal pairing โ€” the foundational associative-learning phenomenon that temporal-difference learning algorithms in RL directly formalize mathematically
๐Ÿ–ฅ๏ธ AI & Computing Sciences DivisionReinforcement Learning & Agents Center