Understanding Knowledge Distillation in Neural Sequence Generation
- Akiko Eriguchi, Jiatao Gu | Microsoft, Facebook AI Research
Sequence-level knowledge distillation (KD) — learning a student model with targets decoded from a pre-trained teacher model — has been widely used in sequence generation applications (e.g. model compression, non-autoregressive translation (NAT), low-resource translation, etc). However, the underlying reasons behind this success have, as of yet, been unclear. In this talk, we will try to tackle the understanding of KD particularly in two scenarios: (1) Learning a weak student from a strong teacher model while keeping the same parallel data used for training the teacher; (2) Learning a student from a teacher model of equal size while the targets are generated from additional monolingual data.
Watch Next
-
Multimodal & Embodied Intelligence (S2)
- Rajiv Ratn Shah,
- Tanuja Ganu,
- Mercy Ranjit
-
Multimodal & Embodied Intelligence (S1), Panel on Multimodal AI: Progress, Pitfalls, Possibilities
- Madhava Krishna,
- Sriram Ganapathy,
- Somak Aditya
-
Session on Compute & Trust (Security)
- Krishna Pillutla,
- Danish Pruthi
-
Panel: Is Retrieval Relevant in the Age of Reasoning?
- Himanshu Tyagi,
- Ravishankar Krishnaswamy,
- Mrinal Kanti Das
-
Session on Reasoning
- Hongxiang Fan,
- Nagarajan Natarajan
-
Human-Centered AI: Design, Deployment & Healthcare
- Manik Gupta,
- Anirudha Joshi,
- Aaditeshwar Seth
-
-
-
-
Beyond Swahili: Designing Inclusive AI for Bantu Languages
- Alfred Malengo Kondoro