Peter Shaw
Staff Research Scientist at Google DeepMind
Google Scholar · Twitter · Bluesky · LinkedIn
Current Work
I work on Gemini post-training, focusing on improving capabilities related to coding, research, and various scientific workflows. My work spans evaluations, data, and infrastructure for agents and reinforcement learning.
I'm also interested in foundational questions related to training recipes and generalization, with implications for safety and reliability. I'm especially motivated by the potential of AI to help address key challenges in biology, medicine, and scientific discovery more broadly.
Published Work
See Google Scholar for a complete list of publications.
Effective Reasoning Chains Reduce Intrinsic Dimensionality
ICML 2026 (Spotlight)
ALTA: Compiler-Based Analysis of Transformers
TMLR 2025 · Code
ProtEx: A Retrieval-Augmented Approach for Protein Function Prediction
MLCB 2025 (Oral) · Code
AgentRewardBench: Evaluating Automatic Evaluations of Web Agent Trajectories
COLM 2025 · Code
Robust Preference Optimization through Reward Model Distillation
TMLR 2025
BAGEL: Bootstrapping Agents by Guiding Exploration with Language
ICML 2024
Helping or Herding? Reward Model Ensembles Mitigate but do not Eliminate Reward Hacking
COLM 2024
From Pixels to UI Actions: Learning to Follow Instructions via Graphical User Interfaces
NeurIPS 2023 (Spotlight) · Code
QUEST: A Retrieval Dataset of Entity-Seeking Queries with Implicit Set Operations
ACL 2023 · Outstanding Paper Award · Code · Dataset
Pix2Struct: Screenshot Parsing as Pretraining for Visual Language Understanding
ICML 2023 · Code
Evaluating the Impact of Model Scale for Compositional Generalization in Semantic Parsing
EMNLP 2022
Improving Compositional Generalization with Latent Structure and Data Augmentation
NAACL 2022 · Code
Compositional Generalization and Natural Language Variation: Can a Semantic Parsing Approach Handle Both?
ACL 2021 · Code
Exploring Unexplored Generalization Challenges for Cross-Database Semantic Parsing
ACL 2020 · Code
Generating Logical Forms from Graph Representations of Text and Entities
ACL 2019
Self-Attention with Relative Position Representations
NAACL 2018 · Code
* Indicates equal contribution.