The main contributions are: (1) a hybrid enhancement-synthesis architecture with selective activation; (2) a localization-guided synthesis mask that restricts synthesis to speech-dominant regions; and (3) an energy-aware training objective that explicitly optimizes the quality-per-compute tradeoff for sustainable edge deployment.
