# MetaXuda: Metal GPU runtime for ML on Apple Silicon (1.1 TOPS with Tokio async)

**URL:** <https://users.rust-lang.org/t/metaxuda-metal-gpu-runtime-for-ml-on-apple-silicon-1-1-tops-with-tokio-async/137649>\
**Category:** announcements\
**Created:** [January 18, 2026, 9:40am UTC](https://users.rust-lang.org/t/metaxuda-metal-gpu-runtime-for-ml-on-apple-silicon-1-1-tops-with-tokio-async/137649 "2026-01-18T09:40:35Z")\
**Posts on this page:** 2\
**Page:** 1

<div class="post-metadata">

**Author:** ![Perinban](https://sea1.discourse-cdn.com/flex019/user_avatar/users.rust-lang.org/perinban/32/51012_2.png) [@Perinban](https://users.rust-lang.org/u/Perinban)\
**Post date:** [January 18, 2026, 9:40am UTC](https://users.rust-lang.org/t/metaxuda-metal-gpu-runtime-for-ml-on-apple-silicon-1-1-tops-with-tokio-async/137649/1 "2026-01-18T09:40:35Z")

</div>

Hey Rustaceans! 👋

I built MetaXuda - a native GPU runtime for machine learning on Apple Silicon, entirely in Rust.

**Motivation:**  
Got tired of "buy Windows for ML" advice. Most ML libraries are CUDA-only with zero macOS GPU support. Translation layers like ZLUDA add overhead, so I built from scratch using Metal.

**Tech Stack:**

- Rust core with Tokio async runtime
- Metal for GPU acceleration
- PyO3 for Python bindings (cuda\_pipeline.so)
- Arrow-based in-kernel quantization
- Multi-tier memory manager (GPU → RAM → SSD)

**Performance:**

- 1.1 TOPS throughput
- 230+ GPU operations (math, transform, ML primitives)
- 93.37% GPU utilization cap (prevents macOS starvation)
- Zero race conditions via centralized scheduler

**Architecture Highlights:**

- Migrated from sync → async (40+ iterations to get it right!)
- Stream managers + thread-pool groups coordinated by scheduler
- Handles 100GB+ workloads through intelligent memory tiering
- CUDA-compatible API naming for library interop

**Current Status:**

- Works with Numba (bypasses execution path)
- pip install metaxuda
- Toolkit integration (scikit-learn, XGBoost) coming next
- CUDA API coverage still in progress

**Known Challenges:**

- Apple's Metal stream limits are undocumented (reverse-engineered what I could)
- Some intentional blocking favors stability over raw speed
- ~1-in-million scheduler notification misses (rare edge case)

**Links:**

- GitHub: [GitHub - Perinban/MetaXuda: A Metal-based CUDA compatibility framework for running Numba CUDA workloads on Apple Silicon. · GitHub](https://github.com/Perinban/MetaXuda-)
- PyPI: pip install metaxuda
- Show HN discussion: [https://news.ycombinator.com/item?id=46664154](https://news.ycombinator.com/item?id=46664154)

**Looking for feedback on:**

- Async scheduler design patterns (Tokio + Metal coordination)
- Memory tier eviction strategies
- Anyone hitting Apple GPU quirks I should know about?

License inquiries: [p.perinban@gmail.com](mailto:p.perinban@gmail.com)

Would love thoughts from the community, especially on the Rust/async architecture choices!

---

<div class="post-metadata">

**Author:** ![system](https://sea1.discourse-cdn.com/flex019/user_avatar/users.rust-lang.org/system/32/51372_2.png) [@system](https://users.rust-lang.org/u/system)\
**Post date:** [April 18, 2026, 9:41am UTC](https://users.rust-lang.org/t/metaxuda-metal-gpu-runtime-for-ml-on-apple-silicon-1-1-tops-with-tokio-async/137649/2 "2026-04-18T09:41:31Z")

</div>

This topic was automatically closed 90 days after the last reply. We invite you to open a new topic if you have further questions or comments.
