Skip to content

juliabloggers.com

A Julia Language Blog Aggregator

  • Home
  • Submit Your RSS/Atom Feed

Matrix multiplication: Performance

By: Ole Kröger

Re-posted from: https://opensourc.es/blog/2021-10-11-matrix-multiplication-performance/index.html

A deep dive into the performance we can obtain by thinking about cache lines and parallel code. An example step by step guide on optimizing dense matrix multiplication.

Related

This entry was posted in Julia and tagged Julia on October 10, 2021 by Ole Kröger.

Post navigation

← Julia beginner’s corner: mastering comparison operators Working with Flux.jl Models on the Hugging Face Hub →

RSS feed

Recent Posts

  • Julia 1.13 Highlights
  • cuTile.jl 1.0: Tile windows, atomics, and sparse views
  • CUDA.jl 6.3: Compiler caching, a new cuDNN, and dependent launches
  • State of Julia's GPU ecosystem in 2026
  • What's New in SciML Since OrdinaryDiffEq v7: Continuation Solvers, Multirate Integrators, and a Pure-Julia Sparse LU
  • OrdinaryDiffEq.jl v7 and DifferentialEquations.jl v8: What Breaks, What to Do About It
  • ModelingToolkit v11: Licensing, Type-Stability / Performance Improvements, and Array Future
  • AMDGPU.jl 2.6 and 2.7: linear algebra, sparse arrays, and RDNA4 matrix cores
  • Metal.jl 1.10: Linear algebra, FFTs, and a faster runtime
  • SciML Small Grants Program: Two Years In, Eight More Projects Funded and Shipped
Proudly powered by WordPress