Foundations of Language Model Security

Theory, Practice, and Fundamental Limits
NeurIPS 2026 Workshop • December 12–13 • Paris, France

About the Workshop

Secure LLM systems by design before insecure patterns become the default.

This workshop aims to advance research on secure-by-design LLM systems by shifting away from the current cat-and-mouse game of attacks and defenses toward a principled understanding of why security vulnerabilities arise and how to address them from the ground up. LLMs have been shown to be vulnerable to a range of attacks such as prompt injections and data poisoning, and yet continue to be deployed in complex systems without a clear understanding of why these vulnerabilities arise or how they interconnect with classical security vulnerabilities.

At the model level, the absence of a hard separation between instructions and data may expose fundamental attack surfaces; at the system level, confused-deputy patterns and missing trust boundaries introduce further structural weaknesses. Understanding whether these vulnerabilities are inherent to current language modeling architectures or artifacts of specific design choices is essential for moving from brittle empirical defenses toward principled security.

Recent research advocates treating LLM security as a system design problem, with a few approaches achieving provable security in specific settings. However, the field still lacks shared formal definitions of what LLM security means, comparable to differential privacy for privacy guarantees, and securing systems by design remains a use-case-specific engineering effort rather than an application of generic principles.

Moreover, existing solutions that offer security guarantees tend to degrade the utility of the system, and it is unclear whether this trade-off is an artifact of current approaches or a more fundamental limit of any LLM system that achieves meaningful security. This workshop aims to consolidate existing knowledge and lay the foundations for future LLM security research by answering three questions:

Q1

Formalizing LLM Security

How should LLM security be formalized? Is there an agreed-upon framework comparable to the notion of differential privacy in privacy research? What role should theory play in creating secure LLM systems?

Q2

Security in Practice

Can we design evaluation methodologies that are reproducible and generalizable rather than fragile and hackable? What concrete steps can help avoid unproductive cycles of attacks and defenses?

Q3

Fundamental Limits of Security

Existing approaches to securing LLM-based systems trade off security for utility. Is this trade-off an artifact of current defense designs, or a more fundamental property that any secure system must exhibit? In which settings has provable security already been achieved, and are there impossibility results establishing conditions under which security cannot be attained?

Workshop Format

The workshop will consist of four thematic blocks. Each block will consist of a 45-minute expert keynote followed by a 30-minute guided discussion, in which participants split into groups of 10-15 with an assigned discussion chair.

The goal of the discussion is to identify open questions that emerge from the keynote. At the end of the discussion, each chair will deliver a 2-minute summary to all attendees. The program will also include a joint poster session, two spotlight contributed talks, and interactive demos to encourage discussion beyond the themes of each block.

45 min
Expert Keynote
30 min
Guided Discussion
10-15
Participants per Group
×4

Invited Speakers

Niloofar Mireshghallah
Niloofar Mireshghallah
humans& and CMU
To be announced
Niloofar is a member of Technical Staff at humans& and incoming Assistant Professor at CMU. She studies privacy, memorization, and contextual integrity in language models.
Christian Schroeder de Witt
University of Oxford
Fundamental Limits of Security
Christian is the Principal Investigator of the Oxford Witt Lab at the University of Oxford. He studies multi-agent security and fundamental limits of assurance in agentic systems.
Reza Shokri
Google and National University of Singapore
Privacy in AI Agents (Q1 and Q2)
Reza is a Senior Staff Research Scientist at Google Zurich and Dean's Chair Associate Professor at NUS. He studies AI security, data privacy, and trustworthy machine learning.
Somesh Jha
University of Wisconsin-Madison
To be announced
Somesh is a Lubar Professor of Computer Sciences at the University of Wisconsin-Madison. He works on security, formal methods, adversarial robustness, and privacy for ML systems.

Demo Presenters

Julia Bazinska
Lakera AI
Interactive demonstration
Julia is a Senior Research Engineer at Lakera AI, where she works on AI security and adversarial evaluation of backbone LLMs in agents.
Matthew Maisel
Sondera
Interactive demonstration
Matthew is the co-founder and CTO of Sondera, where he builds evaluation and runtime policy-enforcement systems for trustworthy AI agents.

Schedule

To be announced

Call for Papers

We invite non-archival submissions of up to 8 pages (excluding references) in the NeurIPS workshop template. We will use OpenReview to manage submissions, as it allows us to streamline notifications and automatically enforce NeurIPS conflict-of-interest policies.

Submissions will be managed through OpenReview. All accepted papers will be presented as posters, with one to two selected for short talks and receiving a best paper award.

Topics of Interest

  • Formal frameworks and definitions for LLM security [Q1]
  • Secure-by-design LLM system architectures [Q1, Q2]
  • Provable security results and impossibility results for LLM-based systems [Q3]
  • Design of reproducible, generalizable evaluation methodologies [Q2]
  • The security–utility trade-off in current and future defenses [Q3]
  • Model- and system-level attacks and defenses (prompt injection, data poisoning, jailbreaks) situated within a principled security framework [Q1, Q2]
  • Multi-agent security and worst-case security guarantees [Q2, Q3]
  • Memorization and privacy in language models [Q1, Q3]
  • Compositional security: do component-level guarantees hold when models or agents are composed into larger systems? [Q3]

Submission Policy

  • In accordance with NeurIPS policy, individuals with a personal conflict of interest with any organizer are not permitted to submit to the workshop.
  • Previously published work is not eligible.
  • We recommend all authors register on OpenReview at least two weeks before the submission deadline.
Submit via OpenReview

Organizers

Egor Zverev
Egor Zverev
Institute of Science and Technology Austria
Maura Pintor
Maura Pintor
University of Cagliari
Santiago Zanella-Beguelin
Santiago Zanella-Béguelin
Microsoft
Ana-Maria Cretu
Ana-Maria Cretu
CISPA Helmholtz Center for Information Security
Nicole Nichols
Nicole Nichols
Palo Alto Networks
Pavel Laskov
Pavel Laskov
University of Liechtenstein

Discussion Chairs

Each thematic block will break into small, moderated groups to identify open questions from the keynote.

To be announced

Supported By

To be announced

Contact

General Questions
Egor Zverev
egor.zverev@ist.ac.at
Submission Questions
Maura Pintor
maura.pintor@unica.it