KeyTalking: Linguistic-Kinematic Keyframe Localization with Prior-Driven Parameter-Efficient Fine-Tuning and Hybrid Inference for Talking Head Generation

Abstract

Existing talking-head generation methods suffer from inference limitations that lead to error accumulation and identity drift in long videos. To address these challenges, we propose KeyTalking, a two-stage framework that integrates linguistic-kinematic keyframe localization with a prior-driven, parameter-efficient fine-tuning procedure and a hybrid inference scheme. First, we detect keyframes from audio via linguistic rules enhanced by privileged kinematic signals and synthesize variable-length keyframe images conditioned on frame-index embeddings and multimodal fusion. Second, we introduce a noise-aware temporal guidance mechanism that injects ReferenceNet features during high-noise steps and employs a PDmotion-Net module during low-noise steps to enforce segment-level start/end constraints and the linear-motion prior. Finally, we use early-stage autoregressive decoding followed by parallel, segment-level refinement during later denoising to achieve long-range coherence with efficient inference. Experiments demonstrate that our fine-tuning and inference strategies significantly reduce computation and that the proposed keyframe localization outperforms end-to-end and naive keyframe baselines in video quality, audio-visual consistency, and long-video stability.

Publication
Proceedings of the SIGGRAPH Asia 2026 Conference Papers (SIGGRAPH Asia 2026)
Shangfei Wang
Shangfei Wang
Professor of Artificial Intelligence

My research interests include Pattern Recognition, Affective Computing, Probabilistic Graphical Models, Computation Intelligence.

Related