Back to Tools Directory

bw24

Rust

memra — from-scratch LLM inference engine for NVIDIA RTX 50-series (Rust + CUDA)

Categories & Topics

aiblackwellcudagemmaggufgpugpu-kernelshy3inferencellama-cppllmmoenvfp4nvidiaperformanceqwenrustsm120aspeculative-decoding
Visit WebsiteView on GitHub

Repository Stats

297
GitHub Stars
35
Forks
297
Watchers
0
Open Issues
License
MIT License
Repository
avifenesh/bw24
Last Updated
Today
August 6, 2026
Latest Release
v0.71.0
Yesterday
Top Contributors
avifenesh
999+
ImgBotApp
1
imgbot[bot]
1
View all contributors →

Related Tools

abliterix

189

Automated alignment adjustment for LLMs — direct steering, LoRA, and MoE expert-granular abliteration, optimized via multi-objective Optuna TPE.

ace-step-ui

4.7k

🎵 The Ultimate Open Source Suno Alternative - Professional UI for ACE-Step 1.5 AI Music Generation. Free, local, unlimited. Stop paying for Suno!

Agent

559

Mac Agent for macOS 26: the agentic AI harness for your Mac Desktop. Computer use, automation, scripting, coding, and more. Powered by 18+ providers across local and cloud LLMs.

agent-native

4.4k

A framework for building agent-native applications.

agent-os

4.3k

Give agents an operating system as a library. Runs in your existing backend – no sandboxes, VMs, or SaaS. Powered by WebAssembly & V8 isolates.

agentmemory

26k

#1 Persistent memory for AI coding agents based on real-world benchmarks