Optimizing Matrix Multiplication on FPGAs using Spatial Parallelism and High-Level Synthesis
Resumo
This paper investigates the optimization of blocked matrix multiplication on the Alveo U50 FPGA using High-Level Synthesis (HLS). The proposed accelerator applies spatial parallelism by through up to four independent Compute Units (CUs) and is evaluated against a multi-threaded CPU implementation. Results show that PCIe transfer overhead limits FPGA performance for small workloads, but scalability becomes advantageous for large matrices. For 4096 × 4096 matrices, the 4-CU design achieves a 4.03× speedup over a single-threaded CPU and matches the performance of an optimized 4-thread CPU implementation while consuming 2.36× less power and dissipating 2.46× less energy. These results highlight the FPGA as a power-efficient solution for large-scale matrix multiplication.Referências
Favaro, F., Dufrechou, E., Oliver, J. P., and Ezzatti, P. (2026). Optimizing high-level synthesis sparse matrix-vector kernel for alveo fpgas. IEEE Transactions on Computer-Aided Design of Integrated Circuits and Systems, pages 1–1.
Gimenez, H. and Midorikawa, E. (2025). Análise comparativa de síntese de alto nível para algoritmos de multiplicação de matrizes em fpga. In Anais da XVI Escola Regional de Alto Desempenho de São Paulo, pages 1–4, Porto Alegre, RS, Brasil. SBC.
Gorlani, P. and Plessl, C. (2021). High level synthesis implementation of a three-dimensional systolic array architecture for matrix multiplications on intel stratix 10 fpgas.
Hosseinabady, M. and Nunez-Yanez, J. L. (2020). A streaming dataflow engine for sparse matrix-vector multiplication using high-level synthesis. IEEE Transactions on Computer-Aided Design of Integrated Circuits and Systems, 39(6):1272–1285.
Kastner, R., Matai, J., and Neuendorffer, S. (2018). Parallel programming for fpgas. CoRR, abs/1805.03648.
Gimenez, H. and Midorikawa, E. (2025). Análise comparativa de síntese de alto nível para algoritmos de multiplicação de matrizes em fpga. In Anais da XVI Escola Regional de Alto Desempenho de São Paulo, pages 1–4, Porto Alegre, RS, Brasil. SBC.
Gorlani, P. and Plessl, C. (2021). High level synthesis implementation of a three-dimensional systolic array architecture for matrix multiplications on intel stratix 10 fpgas.
Hosseinabady, M. and Nunez-Yanez, J. L. (2020). A streaming dataflow engine for sparse matrix-vector multiplication using high-level synthesis. IEEE Transactions on Computer-Aided Design of Integrated Circuits and Systems, 39(6):1272–1285.
Kastner, R., Matai, J., and Neuendorffer, S. (2018). Parallel programming for fpgas. CoRR, abs/1805.03648.
Publicado
02/09/2026
Como Citar
GIMENEZ, Henrique Gregory; MIDORIKAWA, Edson Toshimi; ALMEIDA, Felipe Valencia de; SATO, Liria Matsumoto.
Optimizing Matrix Multiplication on FPGAs using Spatial Parallelism and High-Level Synthesis. In: ESCOLA REGIONAL DE ALTO DESEMPENHO DE SÃO PAULO (ERAD-SP), 17. , 2026, São Paulo/SP.
Anais [...].
Porto Alegre: Sociedade Brasileira de Computação,
2026
.
p. 53-56.
DOI: https://doi.org/10.5753/eradsp.2026.30660.
