Quantization-Triggered Backdoors in Language Models: Cross-Quantizer Transferability and the Validation--Deployment Gap

📡 Tech & Science

Quantization-Triggered Backdoors in Language Models: Cross-Quantizer Transferability and the Validation--Deployment Gap

arXiv:2608.27512v1 Announce Type: new Abstract: Post-training quantization is often treated as a semantically neutral optimization for edge deployment of Large Language Models. When a full-precision source checkpoint is evaluated and quantization is applied downstream without equivalent re-evaluation, this workflow creates a structural validation--deployment gap: because quantization is a many-to-one mapping over parameter space, source-precision certification does not guarantee behavioral equivalence in the deployed configuration. We formalize this gap through Quantization Behavioral Equivalence Classes (QBECs) and prove that QBEC membership does not imply behavioral equivalence, providing a theoretical basis for quantization-triggered backdoor attacks. Building on a three-stage adversarial fine-tuning framework, we embed latent malicious payloads into models that satisfy the source-precision checks used in our evaluation, yet activate targeted adversarial behavior upon INT8 or 4-bit compression. We evaluate this threat in two operationally motivated scenarios, tactical machine translation and political content analysis, extending prior work from decoder-only causal LMs to multilingual encoder-decoder sequence-to-sequence models. Results show that backdoored translation models move from zero measured friend--foe corruption at repaired FP16 to up to 85.02% inversion after quantization, and that a paired stance classifier measures an ideological shift of up to $\Delta\mathrm{Bias}=0.33$ upon compression. A cross-quantizer transferability analysis further shows that attack persistence varies across quantization schemes and model architectures, rather than being determined by nominal bit-width alone. These findings demonstrate that source-precision auditing alone does not rule out quantization-triggered behavior and that the final deployed configuration must be included in behavioral certification for trustworthy edge AI.

📖 Cet article provient d'une source externe.

🔗 Lire l'article complet sur la source →

244 mots extraits · Source originale


🔥 OFFRE PARTENAIRE

Ordinateur portable professionnel ULAP A15ZRC 15,6 pouces, 8 Go de RAM + 512 Go de stockage, Windows 11, processeur AMD Ryzen 5 3500U, prise britannique

🔥 Ordinateur portable professionnel ULAP A15ZRC 15,6 pouces, 8 Go de RAM + 512 Go de stockage, Windows 11, processeur AMD Ryzen 5 3500U, prise britannique - Une offre exceptionnelle à ne pas manquer ! Cliquez pour découvrir.
✅ Consultez les photos supplémentaires.

✅ Découvrez toutes les caractéristiques.

✅ Vérifiez la disponibilité actuelle.

✅ Consultez les avis des acheteurs.

Posts les plus consultés de ce blog

The Best Subscription-Free Home Security Cameras I’ve Tried

Launching Health in ChatGPT

El Niño and Saharan Dust Silence Atlantic Hurricane Season

Accelerating the frontiers of scientific discovery: Google’s $40M commitment to the Genesis Mission

Two Compromised joyfill npm Packages Run RAT When Imported Into Node.js

What Is El Niño? Here’s What It Means for Weather, Water, and Global Economy

OpenAI models used Artifactory zero-days to escape to the internet

Comment supprimer son historique Canal ?

The Anatomy of a $900,000 Validation Bill

Advancing the next era of national science