David Guzman Piedrahita

AI Researcher & ML Engineer · Zurich, Switzerland

David Guzman Piedrahita

ETH Zurich & Max Planck Institute

Research Assistant

 

University of Zürich

MSc Computer Science

I'm an ML researcher and engineer working on AI safety, LLM evaluation, and multi-agent systems. I study how sanctioning norms emerge among self-interested LLM agents in cooperative social dilemmas, and I recently led a position paper on the sociopolitical risks of increasingly agentic AI (ICML 2026 spotlight).

I'm a Research Assistant at ETH Zurich & the Max Planck Institute for Intelligent Systems, supervised by Prof. Zhijing Jin and Bernhard Schölkopf. I'm finishing an MSc in Computer Science at the University of Zürich (AI major, GPA 5.7/6), following a BSc in Computer Science Engineering at the University of Bergamo (107/110, top 1% of faculty). I'm open to opportunities in AI safety research, LLM evaluation, and multi-agent systems.

Email CV Scholar GitHub LinkedIn

Publications

  1. ICML 2026
    Position: Safe Models Do Not Guarantee Safe Societies: The Case for Sociopolitical Risk
    D. Guzman Piedrahita, D. Banerjee, C. Li, T. J. Zhang, K. Blin, S. Simko, P. S. Pandey, I. Strauss, R. Mihalcea, B. Schölkopf, Z. Jin
    Spotlight · Position Paper Track
    Led a position paper arguing that model-level safety does not by itself address the sociopolitical risk posed by increasingly agentic AI systems. The extended preprint, AI Poses Risks to Democratic and Social Systems, was developed with a senior author and advisor network including Yoshua Bengio, Stuart Russell, Audrey Tang, Bernhard Schölkopf, and Richard Mallah.
  2. COLM 2025
    Corrupted by Reasoning: Reasoning Language Models Become Free-Riders in Public Goods Games
    D. Guzman Piedrahita, Y. Yang, M. Sachan, G. Ramponi, B. Schölkopf, Z. Jin
    Best Oral Paper, REALM @ ACL 2025 · 95th-percentile reviews · MSc Thesis
    Investigated the behavior of reasoning-augmented language models in public goods games, finding that reasoning can paradoxically lead to free-riding behavior. Built multi-environment, game-theoretic benchmarks to systematically simulate agent interactions and the influence of punishments and rewards.
  3. EACL 2026
    Democratic or Authoritarian? Probing a New Dimension of Political Biases in Large Language Models
    D. Guzman Piedrahita, I. Strauss, B. Schölkopf, R. Mihalcea, Z. Jin
    Oral presentation
    Proposed a novel methodology to assess LLM alignment with the democracy-authoritarianism spectrum. Found that LLMs generally favor democratic values and leaders, but exhibit increased favorability toward authoritarian figures when prompted in Mandarin.
  4. EMNLP 2025
    Are Language Models Consequentialist or Deontological Moral Reasoners?
    K. Samway, M. Kleiman-Weiner, D. Guzman Piedrahita, R. Mihalcea, B. Schölkopf, Z. Jin
    Introduced a taxonomy of moral rationales to systematically classify reasoning traces by consequentialism and deontology. Revealed that LLM chains-of-thought tend to favor deontological principles, while post-hoc explanations shift toward consequentialist rationales.
  5. Preprint
    When Ethics and Payoffs Diverge: LLM Agents in Morally Charged Social Dilemmas
    S. Backmann, D. Guzman Piedrahita, E. Tewolde, R. Mihalcea, B. Schölkopf, Z. Jin
    Introduced MoralSim (Moral Behavior in Social Dilemma Simulation) to evaluate how LLMs behave in the prisoner's dilemma and public goods game under morally charged contexts. Showed substantial variation across models in both their general tendency to act morally and the consistency of their behavior across game types.
  6. Preprint
    Robustness of Misinformation Classification Systems to Adversarial Examples Through BeamAttack
    A. Fazla, L. Krauter, D. Guzman Piedrahita, A. Michail
    Introduced BeamAttack, a novel algorithm for generating adversarial examples in NLP via beam search. Used a masked language model to predict contextually plausible word replacements, improving the coherence of the generated adversarial examples.
  7. BSc Thesis
    LSTM-based Time Series Forecasting for Air Quality
    D. Guzman Piedrahita
    Evaluated LSTM neural networks for air-quality time-series forecasting in Lombardy, Italy. Employed surrogate models, specifically LIME, to assess feature importance and enhance interpretability.

Experience

Awards & Service

Contact

Open to opportunities in AI safety research, LLM evaluation, and multi-agent systems. Languages: English (TOEFL C1) · Italian (PLIDA C1) · Spanish (native).