DeepSeek 4 Flash and PRO local inference engine for Metal, CUDA and ROCm
Categories & Topics
Repository Stats
Related Tools
bw24
memra — from-scratch LLM inference engine for NVIDIA RTX 50-series (Rust + CUDA)
coreai-torch
Bridges PyTorch and Core AI. Convert existing models to Core AI IR, or author new ones from PyTorch via composite ops, custom op lowerings, and inline Metal GPU kernels.
cuda-oxide
cuda-oxide is an experimental Rust-to-CUDA compiler that lets you write (SIMT) GPU kernels in safe(ish), idiomatic Rust. It compiles standard Rust code directly to PTX — no DSLs, no foreign language bindings, just Rust.
deepseek_v4_rolepaly_instruct
对于DeepSeek-V4角色扮演的特殊控制指令的说明
ds2api
DeepSeek-Compatible Middleware Interface: A technical exploration project in Go, focusing on high-concurrency protocol adaptation. It serves as a reference implementation for converting diverse web protocols into standardized formats.
easy-vibe
💻 The first course for AI-native product builders.