Matrix Multiplication Performance in Multicore Systems with Pthreads, OpenMP, Blocking, SIMD, and BLAS
Abstract
This work evaluates the performance of parallelized matrix multiplication with OpenMP and Pthreads on multicore systems, applying optimizations such as blocking, SIMD (AVX) and use of BLAS. The tests, with matrices of up to 4096×4096 and varying the number of threads, reveal that combinations of these techniques provide significant gains, with speedups above 1000× with BLAS. Pthreads had a slight advantage over OpenMP, and the relevance of blocking and SIMD for cache efficiency and parallelism stands out.
References
Santos, M. S.; Machado, M. C.; Oliveira, D. P. Comparação entre Implementações de Multiplicação de Matrizes com Pthreads, OpenMP e CUDA. Simpósio em Sistemas Computacionais de Alto Desempenho (WSCAD-SSC), 2023. Disponível em: [link]
IEEE. POSIX Threads Programming, 2017. Disponível em: [link]
Intel Corporation. Intel Advanced Vector Extensions (AVX), 2020. Disponível em: [link]
Goto, K., and Van De Geijn, R. Anatomy of High-Performance Matrix Multiplication. ACM Transactions on Mathematical Software, vol. 34, no. 3, pp. 12:1–12:25, 2008.
Ribeiro, B., Carvalho, D., Santos, V. Comparação do uso de OpenMP e Pthreads em uma Paralelização de Multiplicação de Matrizes. ResearchGate, 2019. Disponível em: [link]
Lawson, C.L., Hanson, R.J., Kincaid, D.R., Krogh, F.T. Basic Linear Algebra Subprograms for Fortran Usage. ACM Transactions on Mathematical Software, vol. 5, no. 3, pp. 308-323, 1979. Disponível em: [link]
