David Guzman Piedrahita
PhD Student and Researcher at EPFL · Switzerland
EPFL
Doctoral Assistant · EDIC
I'm a researcher working on AI safety, LLM evaluation, and multi-agent systems. I study how incentives and interactions shape emergent behavior in AI agents, and when aligned behavior holds or drifts under strategic and adversarial pressure. I recently led a position paper on the sociopolitical risks of increasingly agentic AI (ICML 2026 spotlight).
I'm a PhD student in the EDIC doctoral program at EPFL, working with Prof. Robert West and Maksym Andriushchenko at MPI-IS. Previously, I was a Research Assistant at ETH Zurich & the Max Planck Institute for Intelligent Systems, supervised by Prof. Zhijing Jin and Bernhard Schölkopf.
Selected Publications
-
ICML 2026Position: Safe Models Do Not Guarantee Safe Societies: The Case for Sociopolitical RiskSpotlight · Position Paper TrackLed a position paper arguing that model-level safety does not by itself address the sociopolitical risk posed by increasingly agentic AI systems. The extended preprint, AI Poses Risks to Democratic and Social Systems, was developed with a senior author and advisor network including Yoshua Bengio, Stuart Russell, Audrey Tang, Bernhard Schölkopf, and Richard Mallah.
-
COLM 2025Corrupted by Reasoning: Reasoning Language Models Become Free-Riders in Public Goods GamesBest Oral Paper, REALM @ ACL 2025 · 95th-percentile reviews · MSc ThesisInvestigated the behavior of reasoning-augmented language models in public goods games, finding that reasoning can paradoxically lead to free-riding behavior. Built multi-environment, game-theoretic benchmarks to systematically simulate agent interactions and the influence of punishments and rewards.
-
EACL 2026Democratic or Authoritarian? Probing a New Dimension of Political Biases in Large Language ModelsOral presentationProposed a novel methodology to assess LLM alignment with the democracy-authoritarianism spectrum. Found that LLMs generally favor democratic values and leaders, but exhibit increased favorability toward authoritarian figures when prompted in Mandarin.
-
EMNLP 2025Are Language Models Consequentialist or Deontological Moral Reasoners?Introduced a taxonomy of moral rationales to systematically classify reasoning traces by consequentialism and deontology. Revealed that LLM chains-of-thought tend to favor deontological principles, while post-hoc explanations shift toward consequentialist rationales.
-
PreprintWhen Ethics and Payoffs Diverge: LLM Agents in Morally Charged Social DilemmasIntroduced MoralSim (Moral Behavior in Social Dilemma Simulation) to evaluate how LLMs behave in the prisoner's dilemma and public goods game under morally charged contexts. Showed substantial variation across models in both their general tendency to act morally and the consistency of their behavior across game types.
-
PreprintRobustness of Misinformation Classification Systems to Adversarial Examples Through BeamAttackIntroduced BeamAttack, a novel algorithm for generating adversarial examples in NLP via beam search. Used a masked language model to predict contextually plausible word replacements, improving the coherence of the generated adversarial examples.
-
BSc ThesisLSTM-based Time Series Forecasting for Air QualityEvaluated LSTM neural networks for air-quality time-series forecasting in Lombardy, Italy. Employed surrogate models, specifically LIME, to assess feature importance and enhance interpretability.
Experience
- Doctoral Assistant & PhD Student · EPFL, EDIC Sep 2026 – Present
- Research Assistant · ETH Zurich & Max Planck Institute for Intelligent Systems (adv. Z. Jin & B. Schölkopf) Sep 2025 – Aug 2026
- Research Assistant · University of Zürich (LLM fine-tuning, adv. R. Sennrich) Mar – Aug 2024
- Research Assistant · ETH Zurich (large-scale data pipelines) Sep 2023 – Feb 2024
Awards & Service
- ICML 2026 Spotlight · Position Paper Track 2026
- AI Safety Fund Award · co-author, $499,838 2025
- Best Oral Paper · REALM Workshop, ACL 2025 2025
- AI4GOOD Workshop Co-Organizer · ICML 2026 2026
- Oral Presentations · EACL 2026 & IASEAI 2026 (Paris) 2026
- Zurich AI Safety Day · invited poster + $500 Anthropic API credit 2025
- Full Scholarship · University of Bergamo 2020 – 2022
Contact
Interested in collaborations on AI safety, LLM evaluation, and multi-agent systems. Languages: English (TOEFL C1) · Italian (PLIDA C1) · Spanish (native).