Large language models are transforming how researchers analyze text. They can help classify political speeches, identify policy positions, detect misinformation, and process large collections of documents across languages and contexts. Yet they also raise an important question for social science and policy research: when machine-generated classifications are imperfect, how can we ensure that the conclusions drawn from them remain valid? Even when a model has high accuracy, above 90 percent, its remaining errors may not be random and can still bias research findings. To address this challenge, this talk introduces design-based supervised learning, an approach that combines automated annotation with a modest amount of expert human coding. It illustrates the approach with examples from research on online political advertising and censorship in China.
Naoki Egami is an Associate Professor of Political Science at MIT and a Faculty Affiliate of the Institute for Data, Systems, and Society. He specializes in political methodology, with research focusing on causal inference, external validity, machine learning and artificial intelligence, and network and spatial data. His work has appeared in leading political science and computer science journals, including the American Political Science Review, the American Journal of Political Science, and NeurIPS. In 2025, he received the Emerging Scholar Award from the Society for Political Methodology. Egami earned his PhD in Politics from Princeton University and his BA from the University of Tokyo. Before joining MIT, he was an Assistant Professor at Columbia University.
Please RSVP here to sign up for in-person or Zoom attendance. This seminar will be held in E40-496 (Pye Room), and lunch will be available at 11:45am. This event is part of the CIS Global Research & Policy Seminar Series.
Join our mailing list here to learn about upcoming seminars in the series. Contact cis-info@mit.edu with any questions.