David Bau - Publications

Journal Articles

Aaron Mueller, Jannik Brinkmann, Millicent Li, Samuel Marks, Koyena Pal, Nikhil Prakash, Can Rager, Aruna Sankaranarayanan, Arnab Sen Sharma, Jiuding Sun, Eric Todd, David Bau, Yonatan Belinkov. The Quest for the Right Mediator: Surveying Mechanistic Interpretability for NLP Through the Lens of Causal Mediation Analysis. Computational Linguistics 52(1), pp. 331–378 (March 2026).
Lee Sharkey, Bilal Chughtai, Joshua Batson, Jack Lindsey, Jeff Wu, Lucius Bushnaq, Nicholas Goldowsky-Dill, Stefan Heimersheim, Alejandro Ortega, Joseph Bloom, Stella Biderman, Adrià Garriga-Alonso, Arthur Conmy, Neel Nanda, Jessica Rumbelow, Martin Wattenberg, Nandi Schoots, Joseph Miller, William Saunders, Eric J. Michaud, Stephen Casper, Max Tegmark, David Bau, Eric Todd, Atticus Geiger, Mor Geva, Jesse Hoogland, Daniel Murfet, Tom McGrath. Open Problems in Mechanistic Interpretability. Transactions on Machine Learning Research 2025. (TMLR 2025)
Anka Reuel, Benjamin Bucknall, Stephen Casper, Timothy Fist, Lisa Soder, Onni Aarne, Lewis Hammond, Lujain Ibrahim, Alan Chan, Peter Wills, Markus Anderljung, Ben Garfinkel, Lennart Heim, Andrew Trask, Gabriel Mukobi, Rylan Schaeffer, Mauricio Baker, Sara Hooker, Irene Solaiman, Sasha Luccioni, Nitarshan Rajkumar, Nicolas Moës, Jeffrey Ladish, David Bau, Paul Bricman, Neel Guha, Jessica Newman, Yoshua Bengio, Tobin South, Alex Pentland, Sanmi Koyejo, Mykel J. Kochenderfer, Robert Trager. Open Problems in Technical AI Governance. Transactions on Machine Learning Research 2025. (TMLR 2025)
Grace W. Lindsay and David Bau. Testing Methods of Neural Systems Understanding. Cognitive Systems Research 82 (2023), article 101156.
David Bau, Jun-Yan Zhu, Hendrik Strobelt, Agata Lapedriza, Bolei Zhou, and Antonio Torralba. Understanding the Role of Individual Units in a Deep Neural Network. Proceedings of the National Academy of Sciences 117(48), pp. 30071–30078 (2020).
David Bau, Bolei Zhou, Aude Oliva, Antonio Torralba. Interpreting Deep Visual Representations via Network Dissection. IEEE Transactions on Pattern Analysis and Machine Intelligence 41(9), pp. 2131–2145 (2019).
David Bau, Jeff Gray, Caitlin Kelleher, Josh Sheldon, Franklyn Turbak. Learnable Programming: Blocks and Beyond. Communications of the ACM 60(6), pp. 72–80 (2017).

Conference Papers

Kerem Şahin, Sheridan Feucht, Adam Belfki, Jannik Brinkmann, Aaron Mueller, David Bau, Chris Wendler. In-Context Learning Beyond Copying: A Training-Time Analysis of Abstractive ICL. Proceedings of the 2026 Conference on Empirical Methods in Natural Language Processing. (EMNLP 2026). Accepted; forthcoming in 2026.
David Atkinson, Dillon Plunkett, David Bau. Identifying Introspection from the Inside. The 2026 Conference on Language Modeling. (COLM 2026). Accepted; forthcoming in 2026.
Zhuonan Yang, Jacob Xiaochen Li, Francisco Piedrahita Velez, Eric Todd, David Bau, Michael L. Littman, Stephen H. Bach, Ellie Pavlick. Shared Lexical Task Representations Explain Behavioral Variability in LLMs. Proceedings of the 43rd International Conference on Machine Learning, Proceedings of Machine Learning Research 306. (ICML 2026)
Eric Todd, Jannik Brinkmann, Rohit Gandikota, David Bau. In-Context Algebra. Proceedings of the 2026 International Conference on Learning Representations. (ICLR 2026)
Arnab Sen Sharma, Giordano Rogers, Natalie Shapira, David Bau. LLMs Process Lists With General Filter Heads. Proceedings of the 2026 International Conference on Learning Representations. (ICLR 2026)
Nikhil Prakash, Natalie Shapira, Arnab Sen Sharma, Christoph Riedl, Yonatan Belinkov, Tamar Rott Shaham, David Bau, Atticus Geiger. Language Models Use Lookbacks to Track Beliefs. Proceedings of the 2026 International Conference on Learning Representations. (ICLR 2026)
Rohit Gandikota, David Bau. Distilling Diversity and Control in Diffusion Models. IEEE/CVF Winter Conference on Applications of Computer Vision, pp. 1304–1313. (WACV 2026)
Sheridan Feucht, Eric Todd, Byron C. Wallace, David Bau. The Dual-Route Model of Induction. Second Conference on Language Modeling. (COLM 2025)
Rohit Gandikota, Sheridan Feucht, Samuel Marks, David Bau. Erasing Conceptual Knowledge from Language Models. Advances in Neural Information Processing Systems 38, pp. 60681–60713. (NeurIPS 2025)
Viacheslav Surkov, Chris Wendler, Antonio Mari, Mikhail Terekhov, Justin Deschenaux, Robert West, Caglar Gulcehre, David Bau. One-Step is Enough: Sparse Autoencoders for Text-to-Image Diffusion Models. Advances in Neural Information Processing Systems 38, pp. 99956–100027. (NeurIPS 2025)
Kevin Lu, Nicky Kriplani, Rohit Gandikota, Minh Pham, David Bau, Chinmay Hegde, Niv Cohen. When Are Concepts Erased From Diffusion Models? Advances in Neural Information Processing Systems 38, pp. 14751–14775. (NeurIPS 2025)
Jiachen Zhao, Jing Huang, Zhengxuan Wu, David Bau, Weiyan Shi. LLMs Encode Harmfulness and Refusal Separately. Advances in Neural Information Processing Systems 38, pp. 140283–140318. (NeurIPS 2025)
Hui Ren, Joanna Materzyńska, Rohit Gandikota, Giannis Daras, David Bau, Antonio Torralba. Opt-In Art: Learning Art Styles Only from Few Examples. Advances in Neural Information Processing Systems 38, Creative AI Track. (NeurIPS 2025)
A. Feder Cooper, Christopher A. Choquette-Choo, Miranda Bogen, Kevin Klyman, Matthew Jagielski, Katja Filippova, Ken Liu, Alexandra Chouldechova, Jamie Hayes, Yangsibo Huang, Eleni Triantafillou, Peter Kairouz, Nicole Elyse Mitchell, Niloofar Mireshghallah, Abigail Z. Jacobs, James Grimmelmann, Vitaly Shmatikov, Christopher De Sa, Ilia Shumailov, Andreas Terzis, Solon Barocas, Jennifer Wortman Vaughan, Danah Boyd, Yejin Choi, Sanmi Koyejo, Fernando Delgado, Percy Liang, Daniel E. Ho, Pamela Samuelson, Miles Brundage, David Bau, Seth Neel, Hanna Wallach, Amy B. Cyphert, Mark A. Lemley, Nicolas Papernot, Katherine Lee. Machine Unlearning Doesn't Do What You Think: Lessons for Generative AI Policy and Research. Advances in Neural Information Processing Systems 38, Position Paper Track. (NeurIPS 2025)
Rohit Gandikota, Zongze Wu, Richard Zhang, David Bau, Eli Shechtman, Nick Kolkin. SliderSpace: Decomposing the Visual Capabilities of Diffusion Models. Proceedings of the IEEE/CVF International Conference on Computer Vision, pp. 15994–16003. (ICCV 2025)
Aaron Mueller, Atticus Geiger, Sarah Wiegreffe, Dana Arad, Iván Arcuschin, Adam Belfki, Yik Siu Chan, Jaden Fiotto-Kaufman, Tal Haklay, Michael Hanna, Jing Huang, Rohan Gupta, Yaniv Nikankin, Hadas Orgad, Nikhil Prakash, Anja Reusch, Aruna Sankaranarayanan, Shun Shao, Alessandro Stolfo, Martin Tutek, Amir Zur, David Bau, Yonatan Belinkov. MIB: A Mechanistic Interpretability Benchmark. Proceedings of the 42nd International Conference on Machine Learning, Proceedings of Machine Learning Research 267, pp. 45069–45108. (ICML 2025)
Tal Haklay, Hadas Orgad, David Bau, Aaron Mueller, Yonatan Belinkov. Position-aware Automatic Circuit Discovery. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pp. 2792–2817. (ACL 2025)
Hiba Ahsan, Arnab Sen Sharma, Silvio Amir, David Bau, Byron C. Wallace. Elucidating Mechanisms of Demographic Bias in LLMs for Healthcare. Findings of the Association for Computational Linguistics: EMNLP 2025, pp. 14614–14631. (EMNLP Findings 2025)
Jaden Fiotto-Kaufman, Alexander R. Loftus, Eric Todd, Jannik Brinkmann, Koyena Pal, Dmitrii Troitskii, Michael Ripa, Adam Belfki, Can Rager, Caden Juang, Aaron Mueller, Samuel Marks, Arnab Sen Sharma, Francesca Lucchetti, Nikhil Prakash, Carla E. Brodley, Arjun Guha, Jonathan Bell, Byron C. Wallace, David Bau. NNsight and NDIF: Democratizing Access to Open-Weight Foundation Model Internals. Proceedings of the 2025 International Conference on Learning Representations. (ICLR 2025)
Samuel Marks, Can Rager, Eric J Michaud, Yonatan Belinkov, David Bau, Aaron Mueller. Sparse Feature Circuits: Discovering and Editing Interpretable Causal Graphs in Language Models. Proceedings of the 2025 International Conference on Learning Representations. (ICLR 2025 oral)
Koyena Pal, David Bau, Renée J Miller. Model Lakes. Proceedings of the 28th International Conference on Extending Database Technology, pp. 985–995. (EDBT 2025)
Keita Teranishi, Harshitha Menon, William F. Godoy, Prasanna Balaprakash, David Bau, Tal Ben-Nun, Abhinav Bhatele, Franz Franchetti, Michael E. Franusich, Todd Gamblin, Giorgis Georgakoudis, Tom Goldstein, Arjun Guha, Steven E. Hahn, Costin Iancu, Zheming Jin, Terry R. Jones, Tze Meng Low, Het Mankad, Narasinga Rao Miniskar, Mohammad Alaul Haque Monil, Daniel Nichols, Konstantinos Parasyris, Swaroop Pophale, Pedro Valero-Lara, Jeffrey S. Vetter, Samuel Williams, Aaron R. Young. Leveraging AI for Productive and Trustworthy HPC Software: Challenges and Research Directions. Proceedings of the ISC High Performance 2025 International Workshops, pp. 615–625. (ISC 2025)
Imke Grabe, Jaden Fiotto-Kaufman, Rohit Gandikota, David Bau, Tom Jenkins. Hidden Layers: An Interactive Installation for Exploring the Neural Semantics of Image Synthesis. Aarhus Conference 2025 Adjunct Proceedings, article 10, pp. 10:1–10:5. (Aarhus 2025)
Adam Karvonen, Benjamin Wright, Can Rager, Rico Angell, Jannik Brinkmann, Logan Smith, Claudio Mayrink Verdun, David Bau, Samuel Marks. Measuring Progress in Dictionary Learning for Language Model Interpretability with Board Game Models. Advances in Neural Information Processing Systems 37, pp. 83091–83118. (NeurIPS 2024).
Sheridan Feucht, David Atkinson, Byron Wallace, David Bau. Token Erasure as a Footprint of Implicit Vocabulary Items in LLMs. Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing, pp. 9727–9739. (EMNLP 2024)
Arnab Sen Sharma, David Atkinson, David Bau. Locating and Editing Factual Associations in Mamba. Proceedings of the 2024 Conference on Language Modeling. (COLM 2024)
Kenneth Li, Tianle Liu, Naomi Bashkansky, David Bau, Fernanda Viégas, Hanspeter Pfister, Martin Wattenberg. Measuring and Controlling Instruction Instability in Language Model Dialogs. Proceedings of the 2024 Conference on Language Modeling. (COLM 2024)
Stephen Casper, Carson Ezell, Charlotte Siegmann, Noam Kolt, Taylor Lynn Curtis, Benjamin Bucknall, Andreas Haupt, Kevin Wei, Jeremy Scheurer, Marius Hobbhahn, Lee Sharkey, Satyapriya Krishna, Marvin Von Hagen, Silas Alberti, Alan Chan, Qinyi Sun, Michael Gerovitch, David Bau, Max Tegmark, David Krueger, Dylan Hadfield-Menell. Black-Box Access Is Insufficient for Rigorous AI Audits. Proceedings of the 2024 ACM Conference on Fairness, Accountability, and Transparency, pp. 2254–2272. (FAccT 2024)
Evan Hernandez, Arnab Sen Sharma, Tal Haklay, Kevin Meng, Martin Wattenberg, Jacob Andreas, Yonatan Belinkov, David Bau. Linearity of Relation Decoding in Transformer Language Models. Proceedings of the 2024 International Conference on Learning Representations. (ICLR 2024 spotlight)
Eric Todd, Millicent Li, Arnab Sen Sharma, Aaron Mueller, Byron C Wallace, David Bau. Function Vectors in Large Language Models. Proceedings of the 2024 International Conference on Learning Representations. (ICLR 2024)
Nikhil Prakash, Tamar Rott Shaham, Tal Haklay, Yonatan Belinkov, David Bau. Fine-Tuning Enhances Existing Mechanisms: A Case Study on Entity Tracking. Proceedings of the 2024 International Conference on Learning Representations. (ICLR 2024)
Rohit Gandikota, Joanna Materzyńska, Tingrui Zhou, Antonio Torralba, David Bau. Concept Sliders: LoRA Adaptors for Precise Control in Diffusion Models. Proceedings of the European Conference on Computer Vision, pp. 172–188. (ECCV 2024)
Maxwell Jones, Sheng-Yu Wang, Nupur Kumari, David Bau, Jun-Yan Zhu. Customizing Text-to-Image Models with a Single Image Pair. SIGGRAPH Asia 2024 Conference Papers, article 6, pp. 6:1–6:13. (SIGGRAPH Asia 2024)
Rohit Gandikota, Hadas Orgad, Yonatan Belinkov, Joanna Materzyńska, David Bau. Unified Concept Editing in Diffusion Models. Proceedings of the 2024 IEEE/CVF Winter Conference on Applications of Computer Vision, pp. 5099–5108. (WACV 2024)
Koyena Pal, Jiuding Sun, Andrew Yuan, Byron C. Wallace, and David Bau. Future Lens: Anticipating Subsequent Tokens from a Single Hidden State. Proceedings of the 27th Conference on Computational Natural Language Learning, pp. 548–560. (CoNLL 2023)
Sarah Schwettmann, Tamar Rott Shaham, Joanna Materzyńska, Neil Chowdhury, Shuang Li, Jacob Andreas, David Bau, and Antonio Torralba. FIND: A Function Description Benchmark for Evaluating Interpretability Methods. Advances in Neural Information Processing Systems 36, pp. 75688–75715. (NeurIPS 2023).
Rohit Gandikota, Joanna Materzyńska, Jaden Fiotto-Kaufman, David Bau. Erasing Concepts from Diffusion Models. Proceedings of the 2023 IEEE/CVF International Conference on Computer Vision, pp. 2426–2436. (ICCV 2023).
Daohan Lu, Sheng-Yu Wang, Nupur Kumari, Rohan Agarwal, Mia Tang, David Bau, Jun-Yan Zhu. Content-based Search for Deep Generative Models. SIGGRAPH Asia 2023 Conference Papers, article 71, pp. 71:1–71:12. (SIGGRAPH Asia 2023)
Kevin Meng, Arnab Sen Sharma, Alex Andonian, Yonatan Belinkov, David Bau. Mass-Editing Memory in a Transformer. Eleventh International Conference on Learning Representations. (ICLR 2023 spotlight).
Kenneth Li, Aspen K. Hopkins, David Bau, Fernanda Viégas, Hanspeter Pfister, Martin Wattenberg. Emergent World Representations: Exploring a Sequence Model Trained on a Synthetic Task. Eleventh International Conference on Learning Representations. (ICLR 2023 oral).
Kevin Meng, David Bau, Alex Andonian, Yonatan Belinkov. Locating and Editing Factual Associations in GPT. Advances in Neural Information Processing Systems 35, pp. 17359–17372. (NeurIPS 2022).
Sheng-Yu Wang, David Bau, Jun-Yan Zhu. Rewriting Geometric Rules of a GAN. ACM Transactions on Graphics 41(4), article 73, pp. 73:1–73:16. (SIGGRAPH 2022)
Joanna Materzyńska, Antonio Torralba, David Bau. Disentangling Visual and Written Concepts in CLIP. Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp. 16389–16398. (CVPR 2022 oral)
Evan Hernandez, Sarah Schwettmann, David Bau, Teona Bagashvili, Antonio Torralba, Jacob Andreas. Natural Language Descriptions of Deep Visual Features. Proceedings of the International Conference on Learning Representations. (ICLR 2022)
Shibani Santurkar, Dimitris Tsipras, Mahalaxmi Elango, David Bau, Antonio Torralba, and Aleksander Madry. Editing a Classifier by Rewriting Its Prediction Rules. Advances in Neural Information Processing Systems 34, pp. 23359–23373. (NeurIPS 2021)
Emma Andrews, David Bau, and Jeremiah Blanchard. From Droplet to Lilypad: Present and Future of Dual-Modality Environments. 2021 IEEE Symposium on Visual Languages and Human-Centric Computing, pp. 1–2. (VL/HCC 2021)
Sarah Schwettmann, Evan Hernandez, David Bau, Samuel Klein, Jacob Andreas, Antonio Torralba. Toward a Visual Concept Vocabulary for GAN Latent Space. Proceedings of the IEEE/CVF International Conference on Computer Vision, pp. 6804–6812. (ICCV 2021)
Sheng-Yu Wang, David Bau, and Jun-Yan Zhu. Sketch Your Own GAN. Proceedings of the IEEE/CVF International Conference on Computer Vision, pp. 14050–14060. (ICCV 2021)
David Bau, Steven Liu, Tongzhou Wang, Jun-Yan Zhu, and Antonio Torralba. Rewriting a Deep Generative Model. Proceedings of the European Conference on Computer Vision, pp. 351–369. (ECCV 2020 oral)
Lucy Chai, David Bau, Ser-Nam Lim, and Phillip Isola. What Makes Fake Images Detectable? Understanding Properties That Generalize. Proceedings of the European Conference on Computer Vision, pp. 103–120. (ECCV 2020)
Steven Liu, Tongzhou Wang, David Bau, Jun-Yan Zhu, and Antonio Torralba. Diverse Image Generation via Self-Conditioned GANs. Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp. 14274–14283. (CVPR 2020)
David Bau, Jun-Yan Zhu, Jonas Wulff, William Peebles, Hendrik Strobelt, Bolei Zhou, and Antonio Torralba. Seeing What a GAN Cannot Generate. Proceedings of the IEEE/CVF International Conference on Computer Vision, pp. 4502–4511. (ICCV 2019 oral presentation)
David Bau, Hendrik Strobelt, William Peebles, Jonas Wulff, Bolei Zhou, Jun-Yan Zhu, and Antonio Torralba. Semantic Photo Manipulation with a Generative Image Prior. ACM Transactions on Graphics 38(4), article 59, pp. 59:1–59:11. (SIGGRAPH 2019)
Didac Suris, Adria Recasens, David Bau, David Harwath, James Glass, and Antonio Torralba. Learning Words by Drawing Images. Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp. 2029–2038. (CVPR 2019)
David Bau, Jun-Yan Zhu, Hendrik Strobelt, Bolei Zhou, Joshua B. Tenenbaum, William T. Freeman, and Antonio Torralba. GAN Dissection: Visualizing and Understanding Generative Adversarial Networks. Proceedings of the Seventh International Conference on Learning Representations. (ICLR 2019)
David Weintrop, David Bau, and Uri Wilensky. The Cloud Is the Limit: A Case Study of Programming on the Web, with the Web. International Journal of Child-Computer Interaction 20, pp. 1–8. (IJCCI 2019)
Leilani H. Gilpin, David Bau, Ben Z. Yuan, Ayesha Bajwa, Michael Specter, Lalana Kagal. Explaining Explanations: An Overview of Interpretability of Machine Learning. Proceedings of the IEEE 5th International Conference on Data Science and Advanced Analytics, pp. 80–89. (DSAA 2018)
Bolei Zhou, Yiyou Sun, David Bau, and Antonio Torralba. Interpretable Basis Decomposition for Visual Explanation. Proceedings of the European Conference on Computer Vision, pp. 119–134. (ECCV 2018)
David Bau, Bolei Zhou, Aditya Khosla, Aude Oliva, Antonio Torralba. Network Dissection: Quantifying Interpretability of Deep Visual Representations. Proceedings of the 2017 IEEE Conference on Computer Vision and Pattern Recognition, pp. 6541–6549. (CVPR 2017 oral presentation)
David Bau, Matthew Dawson, D. Anthony Bau, C. Sydney Pickens. Pencil Code: Block Code for a Text World. Proceedings of the 14th International Conference on Interaction Design and Children, pp. 445–448. (IDC 2015)
David Bau, Matthew Dawson, D. Anthony Bau. Using Pencil Code to Bridge the Gap Between Visual and Text-Based Coding. Proceedings of the 46th ACM Technical Symposium on Computer Science Education, Abstract Only, p. 706. (SIGCSE 2015)
Ming Zhao, Jay Yagnik, Hartwig Adam, David Bau. Large Scale Learning and Recognition of Faces in Web Videos. 8th IEEE International Conference on Automatic Face and Gesture Recognition, pp. 1–7. (FG 2008)
David Bau, Induprakas Kodukula, Vladimir Kotlyar, Keshav Pingali, Paul Stodghill. Solving Alignment Using Elementary Linear Algebra. Languages and Compilers for Parallel Computing, Lecture Notes in Computer Science 892, pp. 46–60. (LCPC 1994)

Workshop Papers

Rahul Chowdhury, Timothy A. Rupprecht, Senhao Cao, Jiahao Liu, Octavia Camps, David Bau, Pu Zhao, Yanzhi Wang. Mechanistic Interpretability of Structure-Aware Numerical Reasoning in LLaMA 3.1-8B. Proceedings of the 4th Workshop on Rich Media with Generative AI at ACM Multimedia 2026. (RichMediaGAI 2026). Accepted; forthcoming; OpenReview record.
Sheridan Feucht, Byron C Wallace, David Bau. Inducing Induction in Llama via Linear Probe Interventions. The 7th BlackboxNLP Workshop at EMNLP (BlackboxNLP 2024).
Nicholas Vincent, David Bau, Sarah Schwettmann, Joshua Tan. An Alternative to Regulation: The Case for Public AI. Regulatable ML: Towards Bridging the Gaps Between Machine Learning Research and Regulations Workshop at NeurIPS 2023. (RegML 2023).
Silen Naihin, David Atkinson, Marc Green, Merwane Hamadi, Craig Swift, Douglas Schonholtz, Adam Tauman Kalai, David Bau. Testing Language Model Agents Safely in the Wild. Socially Responsible Language Modelling Research workshop at NeurIPS (SoLaR 2023).
Sarah Schwettmann, Neil Chowdhury, Samuel Klein, David Bau, Antonio Torralba. Multimodal Neurons in Pretrained Text-Only Transformers. IEEE/CVF International Conference on Computer Vision Workshops, pp. 2854–2859. (ICCV 2023 Workshop)
Xander Davies, Max Nadeau, Nikhil Prakash, Tamar Rott Shaham, David Bau. Discovering Variable Binding Circuitry with Desiderata. Workshop on Challenges in Deployable Generative AI (ICML 2023 Workshop)
David Bau, Steven Liu, Tongzhou Wang, Jun-Yan Zhu, Antonio Torralba. Horses with Blue Jeans: Creating New Worlds by Rewriting a GAN. 4th Workshop on Machine Learning for Creativity and Design. (NeurIPS 2020 Workshop)
David Bau, Jun-Yan Zhu, Jonas Wulff, William Peebles, Hendrik Strobelt, Bolei Zhou, and Antonio Torralba. Inverting Layers of a Large Generator. ICLR Debugging Machine Learning Models Workshop. (ICLR 2019 workshop)
Jonathan Frankle, David Bau. Dissecting Pruned Neural Networks. ICLR Debugging Machine Learning Models Workshop. (ICLR 2019 workshop)
Saksham Aggarwal, David Anthony Bau, David Bau. A Blocks-Based Editor for HTML Code. IEEE Blocks and Beyond Workshop, pp. 83–85. (VL/HCC 2015 workshop)
David Bau, Anthony Bau. A Preview of Pencil Code: A Tool for Developing Mastery of Programming. Proceedings of the 2nd Workshop on Programming for Mobile & Touch, pp. 21–24. (PROMOTO 2014)

Books

Lloyd N. Trefethen, David Bau. Numerical Linear Algebra. (373pp.) Society for Industrial and Applied Mathematics. (1997)

Submitted Manuscripts and Other Preprints

Complete Manuscripts Under Review

Rohit Gandikota, David Bau. Gaze Heads: How VLMs Look at What They Describe. arXiv:2606.14703 (2026). Submitted; under review at WACV 2027.
Jiachen Zhao, Zhengxuan Wu, Aryaman Arora, Yiyou Sun, David Bau, Weiyan Shi. The Piggyback Hypothesis of Generalization: Explaining and Mitigating Emergent Misalignment. arXiv:2606.06667 (2026). Submitted; under review at NeurIPS 2026.
Kelly Cui, Nikhil Prakash, Shoval Messica, Ayush Raina, David Bau, Antonio Torralba, Tamar Rott Shaham. The Dual Mechanisms of Spatial Variable Binding in Vision–Language Models. arXiv:2603.22278 (2026). Submitted; under review at NeurIPS 2026.
Kevin Lu, Jannik Brinkmann, Stefan Huber, Aaron Mueller, Yonatan Belinkov, David Bau, Chris Wendler. Two Stages of Folding: Convergent Mechanisms in AI Protein Folding Trunks. arXiv:2602.06020 (2026). Submitted; under review at NeurIPS 2026.
Koyena Pal, David Bau, Chandan Singh. Do explanations generalize across large reasoning models? arXiv:2601.11517 (2026). Submitted; under review at NeurIPS 2026.
Can Rager, Chris Wendler, Rohit Gandikota, David Bau. Discovering Forbidden Topics in Language Models. arXiv:2505.17441 (2025). Submitted; under review.

Non-Peer-Reviewed Technical Reports

Natalie Shapira, Chris Wendler, Avery Yen, Gabriele Sarti, Koyena Pal, Olivia Floody, Adam Belfki, Alex Loftus, Aditya Ratan Jannali, Nikhil Prakash, Jasmine Cui, Giordano Rogers, Jannik Brinkmann, Can Rager, Amir Zur, Michael Ripa, Aruna Sankaranarayanan, David Atkinson, Rohit Gandikota, Jaden Fiotto-Kaufman, EunJeong Hwang, Hadas Orgad, P. Sam Sahil, Negev Taglicht, Tomer Shabtay, Atai Ambus, Nitay Alon, Shiri Oron, Ayelet Gordon-Tapiero, Yotam Kaplan, Vered Shwartz, Tamar Rott Shaham, Christoph Riedl, Reuth Mirsky, Maarten Sap, David Manheim, Tomer D. Ullman, David Bau. Agents of Chaos. Research report, arXiv:2602.20021 (2026).
David Bau, Tom McGrath, Sarah Schwettmann, Dylan Hadfield-Menell. AI Dominance Requires Interpretability and Standards for Transparency and Security. Policy white paper submitted in response to the White House OSTP/NITRD request for information on development of an AI Action Plan, March 2025.
Can Rager, David Bau. Auditing AI Bias: The DeepSeek Case. Technical report, January 31, 2025.
Alex Andonian, Sabrina Osmany, Audrey Cui, Yeon-Hwan Park, Ali Jahanian, Antonio Torralba, David Bau. Paint by Word. Technical report, arXiv:2103.10951 (2021).
`