Existing talking-head generation methods suffer from inference limitations that lead to error accumulation and identity drift in long videos. To address these challenges, we propose KeyTalking, a two-stage framework that integrates linguistic-kinematic keyframe localization with a prior-driven, parameter-efficient fine-tuning procedure and a hybrid inference scheme. First, we detect keyframes from audio via linguistic rules enhanced by privileged kinematic signals and synthesize variable-length keyframe images conditioned on frame-index embeddings and multimodal fusion. Second, we introduce a noise-aware temporal guidance mechanism that injects ReferenceNet features during high-noise steps and employs a PDmotion-Net module during low-noise steps to enforce segment-level start/end constraints and the linear-motion prior. Finally, we use early-stage autoregressive decoding followed by parallel, segment-level refinement during later denoising to achieve long-range coherence with efficient inference. Experiments demonstrate that our fine-tuning and inference strategies significantly reduce computation and that the proposed keyframe localization outperforms end-to-end and naive keyframe baselines in video quality, audio-visual consistency, and long-video stability.