Local LLM
Small local LLMs, each given a distinct personality, play an iterated Prisoner's Dilemma tournament against each other
This is a companion project that utilizes Sesame AI's Conversational Speech Model (CSM) in an attempt to recreate the original magic of the companions on release. This is a stack, built and used locally.
This fork of antirez/ds4 is a focused, measured configuration for running DeepSeek V4 Flash on an NVIDIA GeForce RTX 5080 with 16 GB of VRAM, using CUDA and a fast NVMe SSD. Development and performance measurements were made on an RTX 5080 Laptop GPU (Blackwell, sm_120) under Linux/WSL2.
Production-grade recipe for DeepSeek-V4-Flash-0731 (284B MoE) on 2x NVIDIA DGX Spark: self-healing 2-node vLLM cluster, reboot-verified, tuned DSpark speculative decoding (~75 tok/s), full benchmarks, OpenAI Codex CLI integration
Agentic test batteries for local models via Hermes Agent - scripted tools + real Hermes CLI sessions
Run DeepSeek V4 Flash GGUFs on memory-constrained Linux systems using CPU-only NVMe-backed demand paging.
MLX (Apple Silicon) port of Muse-Glimmer-30B. Validated against transformers; also a runtime for the existing MLX conversions, which no released mlx-vlm can load.
All recipes for oss models from Meta Inc.
Apple-style Writing Tools for Hyprland/Omarchy — select text anywhere, hit a keybind, get Translate/Proofread/Correct via Gemini or a local Ollama model.
Art_RAG is a 100% local AI assistant for art history. It analyzes artwork images, builds searchable libraries from PDFs and text files, and offers an interactive chat to query your collection. It features web search fallback, prompting for image generation, and complete privacy.
some tool i decided to make which makes ai type for you in discord
Documentation on small project using an old PC as a home server for storage and running basic LLM.
Redact PII from documents - 100% offline, local LLM, no cloud.
Local multimodal MiniMax H3 prompt writer for ComfyUI, powered by Gemma 4 GGUF models.
NanoMind-S3: Dual-core INT4 inference engine running a 15.2M parameter Stories Transformer (LLaMA-2) on ESP32-S3 @ ~2.96 tok/s. Developed by Salman Farsi (@imfarsi).
Open-source 100% on-device AI voice dictation app for macOS. Free Wispr Flow alternative built with Apple Speech & FoundationModels.