← All projects← 返回项目列表 Efficient ML Algorithm · 2026

Taylor-Calibrate: Principled Initialization for Hybrid Linear Attention Distillation

Lab-led project 实验室主导项目 Converts pretrained Transformers into hybrid GatedDeltaNet students
Taylor-Calibrate framework overview
Abstract 摘要

Taylor-Calibrate is a lightweight initialization method for converting pretrained Transformers into hybrid Gated DeltaNet (GDN) students. Simply copying the teacher's attention projections leaves the new recurrent decay, write, and output-gating dynamics uncalibrated, so the converted model starts in a poor regime and wastes distillation tokens repairing its initialization. Taylor-Calibrate uses Taylor-guided teacher attention statistics to set the value projection, memory timescale, and gates, then applies a short per-layer alignment step. Across four teachers and three retained-layer policies, it produces substantially stronger zero-shot students — up to 88× lower initial perplexity in a representative ablation — and reaches matched recovery targets with 4.9×–9.2× fewer training tokens than naive conversion.

Highlights 亮点 Authors 作者

Zhongzhu Zhou, Qingyang Wu, Junxiong Wang, Mayank Mishra, Shuaiwen Leon Song, Ben Athiwaratkun, Chenfeng Xu

Links 相关链接