Igor Shilov

PhD student at Imperial College London working on AI security & privacy. Big Tech dropout.

prof_pic_new.jpg

about

Hi, I’m Igor!

I’m an AI security & privacy researcher in the final year of my PhD at Imperial College London’s AI Security & Privacy Lab, supervised by Yves-Alexandre de Montjoye.

I was part of the inaugural Anthropic AI Safety Fellowship Program, where I worked on capability removal, and a Research Fellow at Goodfire AI, where I studied mechanistic interpretability.

Before my PhD, I spent over a decade as a software engineer. Most recently, I was at Meta AI as a Research Engineer in Privacy Preserving ML group, where I led the development of Opacus, an open-source PyTorch library for differentially private model training. I also worked on StopNCII.org, a privacy-preserving platform that uses on-device perceptual hashing to combat non-consensual intimate image sharing.

My research interests include:

  • 🏛️ Adversarial robustness in LLMs
  • 🏛️ Memorization in LLMs
  • 🏛️ Model internals

contact

I echo the standing invitation: I like getting email and I like talking to people. Please do reach out if you’re interested in discussing research, collaboration opportunities, or the future of privacy in AI!

I can also be found in Twitter.

selected publications

  1. NeurIPS 2026
    Claudini: Autoresearch discovers state-of-the-art adversarial attack algorithms for LLMs
    Alexander Panfilov*, Peter Romov*, Igor Shilov*, Yves-Alexandre Montjoye, Jonas Geiping, and Maksym Andriushchenko
    In Advances in Neural Information Processing Systems (NeurIPS), 2026
  2. Nature Comms
    The Mosaic Memory of Large Language Models
    Igor Shilov*, Matthieu Meeus*, and Yves-Alexandre Montjoye
    Nature Communications, 2026
  3. arXiv
    Beyond Data Filtering: Knowledge Localization for Capability Removal in LLMs
    Igor Shilov, Alex Cloud, Aryo Pradipta Gema, Jacob Goldman-Wetzler, Nina Panickssery, Henry Sleight, and 2 more authors
    arXiv preprint 2512.05648, 2025
  4. USENIX 2025
    Free Record-Level Privacy Risk Evaluation Through Artifact-Based Methods
    Joseph Pollock*, Igor Shilov*, Euodia Dodd, and Yves-Alexandre Montjoye
    In 34th USENIX Security Symposium (USENIX Security 25), 2025
  5. SaTML 2025
    SoK: Membership Inference Attacks on LLMs are Rushing Nowhere (and How to Fix It)
    Matthieu Meeus, Igor Shilov, Shubham Jain, Manuel Faysse, Marek Rei, and Yves-Alexandre Montjoye
    In 3rd IEEE Conference on Secure and Trustworthy Machine Learning, 2025
  6. ICML 2024
    Copyright Traps for Large Language Models
    Matthieu Meeus*, Igor Shilov*, Manuel Faysse, and Yves-Alexandre Montjoye
    In Forty-first International Conference on Machine Learning (ICML), 2024

    Press coverage in MIT Technology Review and Nature News.

  7. NeurIPS Workshop
    Opacus: User-Friendly Differential Privacy Library in PyTorch
    Ashkan Yousefpour, Igor Shilov, Alexandre Sablayrolles, Davide Testuggine, Karthik Prasad, Mani Malek, and 6 more authors
    In NeurIPS Workshop on Privacy in Machine Learning, 2021

selected projects

  1. opacus2.png
    Opacus: User-Friendly Differential Privacy Library in PyTorch
    Tech Lead (Meta)

    I was the lead developer and maintainer of Opacus, a PyTorch library to train ML models with Differential Privacy. With 1k+ stars on github and 300+ citations, the tool is helping to advance the state-of-the-art in privacy preserving ML both internally and for the wider community of researchers.

  2. stopncii.png
    StopNCII.org: Stop Non-Consensual Intimate Image Abuse
    Tech Lead (Meta)

    I have lead the team developing a privacy-preserving platform helping combat non-consensual intimate image sharing, a joint effort between Meta and a UK-based NGO running “Revenge Porn Helpline”. The platform takes advantage of on-device perceptual hashing to protect privacy. Since it’s launch, many major social media platforms has signed up as industry partners, inclusing Reddit, Snap, TilTok, OnlyFans and PornHub.