Not All 4-bit Quantizers Are Equal: Deployment-Time Mitigation of PII Leakage in Fine-Tuned Small Language Models

Resumo


Organizations fine-tune small language models on private data and then compress them to 4 bits for resource-efficient deployment. We show that which 4-bit method they use also affects privacy, and that what separates the methods is not the bit width but whether they tune their rounding on a small sample of text, the calibration corpus. On our primary model, when each planted record’s own opening text is used as the prompt, the two calibration-based methods we test, Activation-aware Weight Quantization (AWQ) and Gradient-based Post-Training Quantization (GPTQ), each reproduce none of the planted records, while the calibration-corpus-free GGUF Q4 K M format reproduces 5.3% of them on the same seeds. Tracked across five open models with 0.5–7 billion parameters, AWQ leaks least at every size and in both families, with little accuracy loss at 3–7 billion. Controlled experiments associate the difference with calibration-induced rounding error in channels involved in rare-token prediction. Choosing the 4-bit method is therefore a deployment-time privacy decision, not only a question of speed and quality.

Referências

Abitante, J. V. B. et al. (2026). Quantization-robust LLM unlearning via low-rank adaptation. arXiv:2602.13151.

Carlini, N. et al. (2019). The secret sharer: Evaluating and testing unintended memorization in neural networks. In USENIX Security.

Carlini, N. et al. (2021). Extracting training data from large language models. In USENIX Security.

Carlini, N. et al. (2022). Membership inference attacks from first principles. In IEEE S&P.

Carlini, N. et al. (2023). Quantifying memorization across neural language models. In ICLR.

Carlini, N. et al. (2024). Stealing part of a production language model. In ICML.

Clark, P. et al. (2018). Think you have solved question answering? try ARC, the AI2 reasoning challenge. arXiv:1803.05457.

Das, D. et al. (2025). Blind baselines beat membership inference attacks for foundation models. In DATA-FM @ ICLR.

Dettmers, T. et al. (2023). QLoRA: Efficient finetuning of quantized LLMs. In NeurIPS.

Duan, M. et al. (2024). Do membership inference attacks work on large language models? In COLM.

Frantar, E. et al. (2023). OPTQ: Accurate quantization for generative pre-trained transformers. In ICLR.

Geiping, J. et al. (2020). Inverting gradients: How easy is it to break privacy in federated learning? In NeurIPS.

Haque, M. N. et al. (2025). How quantization impacts privacy risk on LLMs for code? arXiv:2508.00128.

Hayes, J. et al. (2025). Measuring memorization in language models via probabilistic extraction. In NAACL.

Husom, E. J. et al. (2025). Sustainable LLM inference for edge AI: Evaluating quantized LLMs for energy efficiency, output accuracy, and inference latency. ACM Transactions on Internet of Things.

Ippolito, D. et al. (2023). Preventing generation of verbatim memorization in language models gives a false sense of privacy. In INLG.

Kandpal, N. et al. (2022). Deduplicating training data mitigates privacy risks in language models. In ICML.

Kapelinski, C. and Kreutz, D. (2026). Decomposing memorization reduction in privacy-preserving fine-tuning of SLMs for CSIRTs. In Brazilian Conference on Intelligent Systems (BRACIS).

Kurt, U. (2026). Which quantization should I use? a unified evaluation of llama.cpp quantization on Llama-3.1-8B-Instruct. arXiv:2601.14277.

Lin, J. et al. (2024). AWQ: Activation-aware weight quantization for on-device LLM compression and acceleration. In MLSys.

Lukas, N. et al. (2023). Analyzing leakage of personally identifiable information in language models. In IEEE S&P.

Maini, P. et al. (2024). LLM dataset inference: Did you train on my dataset? In NeurIPS.

Meeus, M. et al. (2025). SoK: Membership inference attacks on LLMs are rushing nowhere (and how to fix it). In IEEE SaTML.

Mireshghallah, F. et al. (2022). An empirical analysis of memorization in fine-tuned autoregressive language models. In EMNLP.

Mishra, H. and Mehreen, K. (2026). QUAIL: Quantization aware unlearning for mitigating misinformation in LLMs. arXiv:2601.15538.

Nasr, M. et al. (2023). Scalable extraction of training data from (production) language models. arXiv:2311.17035.

Panda, A. et al. (2025). Privacy auditing of large language models. In ICLR.

Sakaguchi, K. et al. (2021). WinoGrande: An adversarial Winograd schema challenge at scale. CACM.

Schwarzschild, A. et al. (2024). Rethinking LLM memorization through the lens of adversarial compression. In NeurIPS.

Shi, W. et al. (2024a). Detecting pretraining data from large language models. In ICLR.

Shi, W. et al. (2024b). MUSE: Machine unlearning six-way evaluation for language models. arXiv:2407.06460.

Wang, F. and Li, B. (2025). Leaner training, lower leakage: Revisiting memorization in LLM fine-tuning with LoRA. arXiv:2506.20856.

Yeom, S. et al. (2018). Privacy risk in machine learning: Analyzing the connection to overfitting. In IEEE CSF.

Zellers, R. et al. (2019). HellaSwag: Can a machine really finish your sentence? In ACL. Zeng, S. et al. (2024). Exploring memorization in fine-tuned language models. In ACL.

Zhang, J. et al. (2025a). Min-k%++: Improved baseline for detecting pre-training data from large language models. In ICLR.

Zhang, Z. et al. (2025b). Catastrophic failure of LLM unlearning via quantization. In ICLR.
Publicado
01/09/2026
KAPELINSKI, Cristhian; KREUTZ, Diego. Not All 4-bit Quantizers Are Equal: Deployment-Time Mitigation of PII Leakage in Fine-Tuned Small Language Models. In: SIMPÓSIO BRASILEIRO DE CIBERSEGURANÇA (SBSEG), 26. , 2026, Armação dos Búzios/RJ. Anais [...]. Porto Alegre: Sociedade Brasileira de Computação, 2026 . p. 833-848. DOI: https://doi.org/10.5753/sbseg.2026.28064.

Artigos mais lidos do(s) mesmo(s) autor(es)

1 2 3 4 5 6 7 8 9 10 > >>