An HLS-Based Hardware Accelerator for ML-KEM Polynomial Operations
Resumo
The rapid advance of quantum computing poses a significant threat to classical cryptographic systems, motivating the NIST standardization of CRYSTALS-Kyber (ML-KEM). Due to the high computational cost of Kyber’s polynomial operations, this paper proposes a High-Level Synthesis (HLS)-based hardware accelerator for Kyber-768, deployed on a PYNQ-Z2 SoC. The proposed architecture adopts a Load-Compute-Store paradigm alongside specific optimization directives to balance resource efficiency and latency. The results are evaluated by comparing the hardware performance to a software-only execution on an ARM Cortex-A9 and to existing state-of-the-art accelerators. The design demonstrates high power efficiency, with a total on-chip power consumption of 1.60W, and low resource utilization, consuming 91 DSP blocks. The highest performance gains occurred in complex vector operations, notably achieving speedups of 4.65× for the Number Theoretic Transform (NTT) and 4.86× for the inverse NTT compared to the software execution. This design suggests that HLS is a viable methodology for accelerating complex post-quantum cryptographic primitives.Referências
Advanced Micro Devices, Inc. (2026). Vitis High-Level Synthesis User Guide (UG1399). AMD, San Jose, CA. v2025.2.
Alagic, G., Dang, Q., Moody, D., Robinson, A., Silberg, H., and Smith-Tone, D. (2024). Module-lattice-based key-encapsulation mechanism standard.
ARM Limited (2016). SoC Designer AXI4 Protocol Bundle User Guide.
Bos, J., Ducas, L., Kiltz, E., Lepoint, T., Lyubashevsky, V., Schanck, J. M., Schwabe, P., Seiler, G., and Stehlé, D. (2018). CRYSTALS-Kyber: A CCA-secure module-lattice-based KEM. In 2018 IEEE European Symposium on Security and Privacy (EuroS&P), pages 353–367. IEEE.
Botros, L., Kannwischer, M. J., and Schwabe, P. (2019). Memory-efficient high-speed implementation of Kyber on Cortex-M4. In International Conference on Cryptology in Africa, pages 209–228. Springer.
Carril, X., Kardaris, C., Ribes-González, J., Farràs, O., Hernandez, C., Kostalabros, V., González-Jiménez, J. U., and Moretó, M. (2024). Hardware acceleration for high-volume operations of CRYSTALS-Kyber and CRYSTALS-Dilithium. ACM Transactions on Reconfigurable Technology and Systems, 17(3).
Chen, H., Chen, H. W., Huang, S. H., and Chen, P. Y. (2025). A high-performance hardware design for polynomial multiplication in the CRYSTALS-Kyber algorithm. In IEEE Symposium on Low-Power and High-Speed Chips and Systems, COOL CHIPS 2025 - Proceedings. Institute of Electrical and Electronics Engineers Inc.
Kastner, R., Matai, J., and Neuendorffer, S. (2018). Parallel Programming for FPGAs. ArXiv e-prints.
Regev, O. (2005). On lattices, learning with errors, random linear codes, and cryptography. In Proceedings of the Thirty-Seventh Annual ACM Symposium on Theory of Computing, STOC ’05, page 84–93, New York, NY, USA. Association for Computing Machinery.
Shor, P. W. (1997). Polynomial-time algorithms for prime factorization and discrete logarithms on a quantum computer. SIAM Journal on Computing, 26(5):1484–1509.
Soni, D. and Karri, R. (2021). Efficient hardware implementation of PQC primitives and PQC algorithms using high-level synthesis. In Proceedings of IEEE Computer Society Annual Symposium on VLSI, ISVLSI, volume 2021-July, pages 296–301. IEEE Computer Society.
Sun, J. and Bai, X. (2024). A high-speed hardware architecture of an NTT accelerator for CRYSTALS-Kyber. Integrated Circuits and Systems, 1:92–102.
TUL Corporation (2018). PYNQ-Z2 Reference Manual. TUL Corporation. v1.0.
Véstias, M., Flores, P., and Cláudio de Campos Neto, H. (2025). Síntese de alto nível em fpga.
Yaman, F., Mert, A. C., Öztürk, E., and Savaş, E. (2021). A Hardware Accelerator for Polynomial Multiplication Operation of CRYSTALS-Kyber PQC Scheme. IEEE.
Zhang, Z., Cui, Y., Ni, Z., Wang, C., and Liu, W. (2023). An efficient hardware accelerator of high-speed NTT for CRYSTALS-Kyber post-quantum cryptography. In Conference Record - Asilomar Conference on Signals, Systems and Computers, pages 1–6. IEEE Computer Society.
Alagic, G., Dang, Q., Moody, D., Robinson, A., Silberg, H., and Smith-Tone, D. (2024). Module-lattice-based key-encapsulation mechanism standard.
ARM Limited (2016). SoC Designer AXI4 Protocol Bundle User Guide.
Bos, J., Ducas, L., Kiltz, E., Lepoint, T., Lyubashevsky, V., Schanck, J. M., Schwabe, P., Seiler, G., and Stehlé, D. (2018). CRYSTALS-Kyber: A CCA-secure module-lattice-based KEM. In 2018 IEEE European Symposium on Security and Privacy (EuroS&P), pages 353–367. IEEE.
Botros, L., Kannwischer, M. J., and Schwabe, P. (2019). Memory-efficient high-speed implementation of Kyber on Cortex-M4. In International Conference on Cryptology in Africa, pages 209–228. Springer.
Carril, X., Kardaris, C., Ribes-González, J., Farràs, O., Hernandez, C., Kostalabros, V., González-Jiménez, J. U., and Moretó, M. (2024). Hardware acceleration for high-volume operations of CRYSTALS-Kyber and CRYSTALS-Dilithium. ACM Transactions on Reconfigurable Technology and Systems, 17(3).
Chen, H., Chen, H. W., Huang, S. H., and Chen, P. Y. (2025). A high-performance hardware design for polynomial multiplication in the CRYSTALS-Kyber algorithm. In IEEE Symposium on Low-Power and High-Speed Chips and Systems, COOL CHIPS 2025 - Proceedings. Institute of Electrical and Electronics Engineers Inc.
Kastner, R., Matai, J., and Neuendorffer, S. (2018). Parallel Programming for FPGAs. ArXiv e-prints.
Regev, O. (2005). On lattices, learning with errors, random linear codes, and cryptography. In Proceedings of the Thirty-Seventh Annual ACM Symposium on Theory of Computing, STOC ’05, page 84–93, New York, NY, USA. Association for Computing Machinery.
Shor, P. W. (1997). Polynomial-time algorithms for prime factorization and discrete logarithms on a quantum computer. SIAM Journal on Computing, 26(5):1484–1509.
Soni, D. and Karri, R. (2021). Efficient hardware implementation of PQC primitives and PQC algorithms using high-level synthesis. In Proceedings of IEEE Computer Society Annual Symposium on VLSI, ISVLSI, volume 2021-July, pages 296–301. IEEE Computer Society.
Sun, J. and Bai, X. (2024). A high-speed hardware architecture of an NTT accelerator for CRYSTALS-Kyber. Integrated Circuits and Systems, 1:92–102.
TUL Corporation (2018). PYNQ-Z2 Reference Manual. TUL Corporation. v1.0.
Véstias, M., Flores, P., and Cláudio de Campos Neto, H. (2025). Síntese de alto nível em fpga.
Yaman, F., Mert, A. C., Öztürk, E., and Savaş, E. (2021). A Hardware Accelerator for Polynomial Multiplication Operation of CRYSTALS-Kyber PQC Scheme. IEEE.
Zhang, Z., Cui, Y., Ni, Z., Wang, C., and Liu, W. (2023). An efficient hardware accelerator of high-speed NTT for CRYSTALS-Kyber post-quantum cryptography. In Conference Record - Asilomar Conference on Signals, Systems and Computers, pages 1–6. IEEE Computer Society.
Publicado
01/09/2026
Como Citar
GIMENEZ, Henrique Gregory; MIDORIKAWA, Edson Toshimi; ALMEIDA, Felipe Valência de; PAIVA, Thales B..
An HLS-Based Hardware Accelerator for ML-KEM Polynomial Operations. In: WORKSHOP DE TRABALHOS DE INICIAÇÃO CIENTÍFICA E DE GRADUAÇÃO - SIMPÓSIO BRASILEIRO DE CIBERSEGURANÇA (SBSEG), 26. , 2026, Armação dos Búzios/RJ.
Anais [...].
Porto Alegre: Sociedade Brasileira de Computação,
2026
.
p. 358-368.
DOI: https://doi.org/10.5753/sbseg_estendido.2026.29274.
