CS 294-288: Data-Centric LLMs
Fall 2026
Instructor: Sewon Min
Class hours: TuThu 14:00-15:30 (14:10-15:30 considering Berkeley time)
Class location: Gateway B1023
Office hours: By appointment
Contact: Slack DM (we use Slack for all course communication)
Overview: Advances in large language models (LLMs) have been driven by the increasing availability of large, diverse datasets. But where do these datasets come from, how are they used, and how can we leverage them more effectively? This course explores these questions, examining what data we use, how and why it works, and the challenges it introduces in LLM development.
The course is primarily designed for PhD students and centers on paper readings, discussions, and an open-ended project. Students are expected to have a strong background in ML/NLP/LLMs and be familiar with CS 288 materials, with the ability to independently engage with research papers.
Class Syllabus (Tentative)
All deadlines are at 5:59 PM PT.
- 08/27 Thu
- Introduction [Slides]
- 09/01 Tue
- Pre-training data curation
- Language Models are Few-Shot Learners
- DataComp-LM: In search of the next generation of training sets for language models
- FineWeb: decanting the web for the finest text data at scale
- Additional readings
- Language Models are Unsupervised Multitask Learners (Sec 2.1)
- Exploring the Limits of Transfer Learning with a Unified Text-to-Text Transformer (Sec 2.2)
- Deduplicating Training Data Makes Language Models Better
- The Pile: An 800GB Dataset of Diverse Text for Language Modeling
- Dolma: an Open Corpus of Three Trillion Tokens for Language Model Pretraining Research
- Nemotron-CC: Transforming Common Crawl into a Refined Long-Horizon Pretraining Dataset
- 09/03 Thu
- Guest lecture by Shayne Longpre (MIT PhD, Anthropic)
- Talk title: TBA
- 09/08 Tue
- Scaling laws
- Prerequisite
- Main readings
- Language models scale reliably with over-training and on downstream tasks
- Evaluating the Robustness of Chinchilla Compute-Optimal Scaling
- Small-Scale Experiments: Are We There Yet?
- Recommended optional readings: MoE scaling laws
- Prerequisite
- 09/10 Thu
- Infinite compute scaling laws
- Prerequisite
- Scaling Data-Constrained Language Models
- Scaling Laws for Data Filtering – Data Curation cannot be Compute Agnostic
- Main readings
- Additional readings
- Prerequisite
- 09/15 Tue
- Data provenance and representation
- 09/17 Thu
- Data copyright and permissivity
- Foundation Models and Fair Use
- Consent in Crisis: The Rapid Decline of the AI Data Commons
- The Common Pile v0.1: An 8TB Dataset of Public Domain and Openly Licensed Text
- Additional readings
- 09/22 Tue
- Synthetic pre-training
- Prerequisite
- Textbooks are all you need
- Cosmopedia: how to create large-scale synthetic data for pre-training
- Synthetic pretraining
- Main readings
- Prerequisite
- 09/24 Thu
- Model collapse and the future data ecosystem
- 09/29 Tue
- Will we really run out of data?
- 10/01 Thu
- Special topic: We will choose either Option A or Option B
- Option A: Frontier Open-Source LLM
- If we choose Option A, we will select one paper from the following list.
- The Llama 3 Herd of Models (2024)
- DeepSeek-V3 Technical Report (2024)
- DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning (2025)
- GLM-5: from Vibe Coding to Agentic Engineering (2026)
- DeepSeek-V4 Technical Report (2026)
- Option B: Next-Generation Architecture
- If we choose Option B, we will select two papers from the following list.
- Option A: Frontier Open-Source LLM
- 10/06 Tue
- No class: Replacing it with offline feedback sessions
- 10/08 Thu
- No class: Replacing it with offline feedback sessions
- 10/13 Tue
- Class activity: Discussion of talks from the BAIR-NLP Workshop
- The BAIR-NLP Workshop is an all-day event on Monday, October 5.
- 10/15 Thu
- Midpoint presentations
- 10/20 Tue
- Midpoint presentations
- Project midpoint report due
- 10/22 Thu
- Guest lecture (TBA)
- 10/27 Tue
- AI watermarking
- Additional readings
- 10/29 Thu
- AI generated text detection
- Main readings
- Additional readings
- Main readings
- 11/03 Tue
- Creativity, copying, and homogenization
- Prerequisite
- Main readings
- Death of the Novel(ty): Beyond n-Gram Novelty as a Metric for Textual Creativity
- Can Good Writing Be Generative? Expert-Level AI Writing Emerges through Fine-Tuning on High-Quality Books
- Additional readings
- How much do language models copy from their training data? Evaluating linguistic novelty in text generation using RAVEN
- AI as Humanity’s Salieri: Quantifying Linguistic Creativity of Language Models via Systematic Attribution of Machine Text against Web Text
- Readers Prefer Outputs of AI Trained on Copyrighted Books over Expert Human Writers
- Measuring AI “Slop” in Text
- Generative AI floods and dilutes the market for books
- Prerequisite
- 11/05 Thu
- Training data attribution and valuation
- How can we distinguish copying, causal influence, and economic value?
- Prerequisite
- TRAK: Attributing Model Behavior at Scale
- Studying Large Language Model Generalization with Influence Functions
- OLMoTrace: Tracing Language Model Outputs Back to Trillions of Training Tokens
- Main readings
- Scalable Influence and Fact Tracing for Large Language Model Pretraining
- What is Your Data Worth to GPT? LLM-Scale Data Valuation with Influence Functions
- Additional readings
- How can we distinguish copying, causal influence, and economic value?
- 11/10 Tue
- Choose one of Membership inference and Training data extraction
- Option 1: Membership inference
- Prerequisite
- Detecting Pretraining Data from Large Language Models
- Do Membership Inference Attacks Work on Large Language Models?
- LLM Dataset Inference: Did you train on my dataset?
- Main readings
- Reassessing EMNLP 2024’s Best Paper: Does Divergence-Based Calibration for Membership Inference Attacks Hold Up?
- Membership Inference Attacks Cannot Prove that a Model Was Trained On Your Data
- Option 2: Training data extraction
- Prerequisite
- Main readings
- Language Models May Verbatim Complete Text They Were Not Explicitly Trained On
- Extracting memorized pieces of (copyrighted) books from open-weight language models
- Extracting books from production language models
- Measuring memorization in language models via probabilistic extraction
- Additional readings
- Option 1: Membership inference
- 11/12 Thu
- Final presentations
- 11/17 Tue
- Final presentations
- 11/19 Thu
- Final presentations
- 11/24 Tue
- Final presentations
- 11/26 Thu
- No class: Thanksgiving
- 12/01 Tue
- No class: Replacing it with offline feedback sessions
- 12/03 Thu
- No class: Replacing it with offline feedback sessions
- Project final report due by 12/14 (Mon)