BlackRiver / Local intelligence

Intelligence,under your control.

Browser-native products, native llama.cpp releases and model-engineering research built to keep inference close to the hardware. From compact WebGPU models to 284B-class GGUF systems.

08Public systems
2Runtime classes
0Server inference

Choose a system.

Each lab is a working product surface or a documented model release. Browser labs require WebGPU; native releases target local runtimes such as llama.cpp.

Chat / Documents / Vision / Audio / Local OCR
Production workspaceOpen Workspace →

BlackRiver Workspace

A complete private AI workspace for conversations, documents, scanned-PDF OCR, images and audio. Choose Gemma 4 E2B or E4B and run the full experience locally in your browser.

Powered by the proven BlackRiver Gemma 4 WebGPU runtime and local document pipeline.BlackRiver AI product: model management, chat, local storage, OCR, citations, multimodal attachments and resilient browser inference.

Production V1E2B / E4BDocumentsMultimodalPrivate
Qwen3.6 foundation / BlackRiver Phase 10.2 post-training / MTP-complete GGUF
Flagship local modelExplore release →

QwiVer3.6-35B-A3B

BlackRiver's daily-driver sparse MoE for coding, agents and long-horizon reasoning. In our internal real-world workflow evaluation it consistently beats the upstream Qwen3.6-35B-A3B we started from.

Base model: Qwen/Qwen3.6-35B-A3B · Training base: unsloth/Qwen3.6-35B-A3B.Created by A.I Joe and published by BlackRiver AI Ltd: Phase 10.2 curriculum post-training, exact BF16 adapter merge, MTP-preserving GGUF conversion, multimodal projector and release validation.

35B / ~3B active262K nativeVisionNative MTPQ2 / Q3 / Q4 / Q8Public release
DeepSeek V4 Flash 0731 / Quantization-aware behavioral modification
Open-weight model researchExplore release →

DeepRiver V4 Flash

Two llama.cpp-ready GGUF variants engineered to preserve capability while substantially reducing refusal behavior. Flash prioritizes maximum practical context; Flash Pro uses a higher-precision IQ3_XXS foundation.

Base model: DeepSeek V4 Flash 0731 · Quantized foundations: Unsloth UD-Q2_K_XL and UD-IQ3_XXS.Created by A.I Joe and published by BlackRiver AI Ltd: refusal-direction projection, norm-preserved Q8_0 payloads, byte-range patching and full integrity audit.

284.3B classFlash / Prollama.cppGGUF≈14 tok/s testedRelease candidate
Multilingual TTS / Voice cloning / Automatic ASR
Voice intelligenceLaunch →

OmniVoice

Generate multilingual speech and clone a reference voice entirely in the browser. Parakeet transcribes the reference automatically before private, local WebGPU synthesis.

Core TTS: OmniVoice by k2-fsa · Automatic speech recognition: NVIDIA Parakeet · Browser ONNX packaging and community implementation: VocoLoco.BlackRiver AI implementation: unified loading flow, automatic transcription, long-form generation, interface and deployment.

646 languagesVoice cloningParakeet ASRNo usage quotaWebGPU
27B / 1-bit
Language modelLaunch →

Bonsai 27B

Large-scale private reasoning with a dense 27-billion-parameter model compressed to 1-bit weights.

Model and low-bit research: Prism ML · Base architecture: Qwen3.6-27B · Browser foundation: WebML Community.BlackRiver AI implementation: product interface, integration and deployment.

Text3.8 GBWebGPU
Vision / Audio / Text
Multimodal modelLaunch →

Gemma 4

Understand images, transcribe audio and converse with a compact multimodal model entirely in your browser.

Model: Google · Browser foundation: WebML Community and Hugging Face Transformers.js · WebGPU kernels: Fable 5.BlackRiver AI implementation: multimodal adaptation, interface and site integration.

E2B / E4BVisionAudio
High-speed inference
Compact language modelLaunch →

LFM2.5 230M

A compact Liquid AI language model tuned for extremely fast local generation and structured tasks.

Model: Liquid AI · Browser foundation: WebML Community · Original kernel optimization credits: Fable 5 and Opus 4.8.BlackRiver AI implementation: product interface, navigation and deployment.

230MTextQ4_0
Local diffusion
Image modelLaunch →

Bonsai Image

Generate high-quality images locally with compressed 4B diffusion models and private prompts.

Low-bit models: Prism ML · Base model: FLUX.2 Klein 4B by Black Forest Labs · Browser foundation: WebML Community.BlackRiver AI implementation: product interface, integration and deployment.

4BImageBinary / Ternary

Credits,
clearly.

“BlackRiver AI implementation” describes the unified product interface, navigation, deployment packaging and selected workflow adaptations for third-party browser research. QwiVer and DeepRiver are separately identified as BlackRiver model-engineering releases built from fully credited upstream foundations.

QwiVer3.6-35B-A3B

Qwen · Unsloth · A.I Joe · BlackRiver AI

QwiVer is derived from Qwen3.6-35B-A3B and retains upstream Qwen architecture credit. A.I Joe / BlackRiver AI developed the Phase 10.2 curriculum post-training, exact BF16 adapter merge, MTP-preserving GGUF release pipeline, multimodal packaging and validation. Read the complete QwiVer presentation.

DeepRiver V4 Flash

DeepSeek · Unsloth · A.I Joe · BlackRiver AI

DeepRiver is derived from DeepSeek V4 Flash 0731 and preserves the original Unsloth quantization credit. A.I Joe / BlackRiver AI developed the quantization-aware behavioral modification, direct Q8_0 replacement payloads, GGUF range-patching pipeline, integrity manifests and release packaging. Read the complete technical presentation.

OmniVoice

k2-fsa · NVIDIA · VocoLoco

k2-fsa created OmniVoice. Automatic reference transcription uses NVIDIA Parakeet. The browser-ready ONNX work and model packaging are credited to VocoLoco. BlackRiver AI integrated the complete local workflow, interface and deployment surface.

Bonsai 27B

Prism ML · Qwen · WebML Community

Prism ML created the Bonsai 27B low-bit model, derived from Qwen3.6-27B. The original browser work is credited to WebML Community.

Gemma 4 Multimodal

Google · Hugging Face · Fable 5

Google created Gemma 4. The browser foundation comes from WebML Community and Hugging Face tooling; the original optimized WebGPU kernels are credited to Fable 5.

LFM2.5 230M

Liquid AI · WebML Community

Liquid AI created LFM2.5 230M. The original browser demo is credited to WebML Community, with kernel optimization credits retained from that project.

Private by architecture.

Inference happens on your GPU. Content does not need to leave the browser.

Real model weights.

Every runtime uses real quantized weights and documented foundations, with no simulated outputs.

Built for local compute.

WebGPU makes browser AI accessible; llama.cpp brings frontier-scale GGUF models to native hardware.