We are a group for people interested in running, evaluating, and building with AI locally. We are not interested in how GPU Rich you are -- we are a community focused on digging into relevant research, tools, and practical techniques for hosting models on consumer and prosumer hardware. We care about comparing real-world performance across different systems and models and sharing local agentic workflows we actually use day to day -- in as few FLOPs as you can muster.
Topics will range from inference stacks like vLLM, SGLang, and llama.cpp to hardware tradeoffs, model quality, optimization, benchmarking, and emerging local-first AI tooling. The scope is intentionally broad: if it helps us better understand, run, or build useful AI systems without relying entirely on the cloud, it’s fair game.
Join us and free the FLOPs.
Cohere Labs Discord Channel: #local-ai
Upcoming Session