The Universal AI Tactical Workbench & Custom Model Foundry.
Connect any frontier model—OpenAI, Anthropic, Gemini, Mistral—or run curated local models and train your own custom LoRAs. One unified desktop command center for generation, synthetic training data, fine-tuning, and live deployment.
Universal Model Freedom
Zero vendor lock-in. Switch effortlessly between frontier cloud intelligence and private offline weights with unified prompt contracts, dynamic token counters, and microsecond telemetry.
Run Sovereign Weights Locally. Zero Cloud Costs.
Download pre-quantized, hardware-optimized local checkpoints directly inside Howitzer. Fully air-gapped with Metal GPU offloading.
Qwen 2.5 Coder
State-of-the-art code generation, refactoring, and AST analysis. Optimized for sub-15ms time-to-first-token on Apple Silicon.
DeepSeek R1 Distill
Dense chain-of-thought verification, algorithmic deduction, and logic proofs distilled into an efficient local footprint.
Llama 3.3 Instruct
High-fidelity general instruction following, dynamic schema extraction, and strict JSON output formatting on device.
Mistral NeMo 12B
NVIDIA & Mistral co-designed multilingual foundation model. High-capacity reasoning for long-context research and dossiers.
The Foundry: Train, Quantize & Deploy
Howitzer isn't just an API consumer. With the integrated Foundry studio, you can autonomously synthesize training data, fine-tune custom LoRA adapters directly on your Mac, quantize to GGUF, and evaluate live with zero cloud dependencies.
Agentic Synthetic Data Studio
Ingest raw seed documentation, code repositories, or PDFs. Autonomous multi-agent pipelines generate thousands of instruction-response pairs with automated critique and rejection sampling.
- Sampling:Adversarial Rejection Passes
- Export:JSONL, Parquet, Alpaca, ChatML
- Deduplication:MinHash LSH & Semantic AST Filter
Desktop LoRA on Apple Silicon
Fine-tune 3B to 8B models directly on Apple Silicon Metal GPUs. Monitor live loss curves, learning rate warmups, and perplexity meters in real time without renting expensive cloud clusters.
- Acceleration:Apple Silicon Metal Performance Shaders
- Architecture:LoRA & QLoRA (Rank 8..64)
- Hot-Swapping:Dynamic Adapter Swapping at Runtime
1-Click GGUF & MLX Quantization
Package your fine-tuned weights for immediate edge or desktop distribution. Merge LoRA weights into base models and quantize down to Q4_K_M or Q8_0 with a single keystroke.
- Target Formats:GGUF, MLX, SafeTensors
- Perplexity Check:Automated Degradation Guard
- Air-Gap:Standalone Portable Model Bundles
Live Canary & Regression Range
Test your fine-tuned adapters side-by-side against frontier baselines and the original base model. Run automated regression benchmarks to ensure zero catastrophic forgetting before deployment.
- Evaluation:Automated Regression Assertion Suites
- Side-by-Side:Base Model vs. Custom LoRA vs. Frontier
- Local Serving:Instant OpenAI-Compatible HTTP Server
The Five Execution Batteries
Engineered with pure Qt 6 C++ on Apple Silicon. Every control, table, slider, and lanyard operates within deterministic in-memory state under the Blunt UI paradigm.
DirectFire
Deterministic API workbench. Inspect request headers, route query params, manage auth tokens, configure payloads, and analyze instantaneous response telemetry.
- Inspector:Status / Latency / Payload Badges
- Protocol:HTTP 1.1 / HTTP 2 / SSL Pinning
- Trigger:[ PULL ] Tactical Action Button
Arsenal
Dynamic prompt engineering studio. Substitute template variables via mustache bracket syntax, dial temperature and top-p hyperparameters, and enforce strict structured JSON schemas.
- Substitution:Dynamic Variable Matrix
- Schema Mode:Strict JSON AST Validator
- Trigger:[ PULL ARSENAL ] Lanyard
Battery
Chained workflow pipeline editor. Compose multi-stage execution sequences (Stages 1..N), configure regex and JSONPath extraction rules, and dispatch in parallel or sequentially.
- Sequencer:Reorderable Stage Tree (1..N)
- Execution:Sequential & Parallel Dispatch
- Log Stream:Live Terminal Telemetry Console
Firing Range
Tri-model shootout and diff arena. Pit Gemini 3.5 Flash, Local Qwen 2.5 8B, and Claude 3.5 Sonnet side-by-side. Inspect TTFT, latency, token costs, and trigger automated arbiter synthesis.
- Matrix:3-Column Parallel Comparison
- Diff Analysis:Unified Character & Token Diff
- Arbiter:Synthesis & Consensus Generator
Magazine
Secure hardware enclave monitor and environment secret vault. Toggle hardware security simulation, switch environments (Prod/Staging/Local), and export shell configurations via eval hooks.
- Vault:Masked / Revealed Secret State
- Enclave:Hardware Simulation Guard
- Shell Hook:eval $(howitzer env ...)
Predictable Pricing for Sovereign Builders
Local-first freemium. Run unlimited local models and API tests for zero dollars forever. Upgrade to Pro for the Foundry creation studio and cloud shootouts.
Permanent local AI workbench for developers, students, and open-source engineers.
- ✓ Unlimited Local GGUF Inference (Qwen, Llama)
- ✓ DirectFire API & HTTP Workbench
- ✓ Arsenal Prompt Studio & Template Variables
- ✓ Bring Your Own Frontier Keys (OpenAI, Claude, Gemini)
- ✓ AES-256 Air-Gapped Keychain Storage
- ✓ Howitzer CLI Basic Shell Integration
Complete model creation suite, synthetic dataset generation, and desktop LoRA fine-tuning.
- ✓ Everything in Community Tier
- ✓ The Foundry Studio (Synthetic Data Generator)
- ✓ Desktop LoRA Tuning on Apple Silicon Metal
- ✓ Live Canary Range & Automated Regression Evals
- ✓ Firing Range 3-Model Shootout & Arbiter Synthesis
- ✓ Battery Multi-Stage Pipeline Sequencer
- ✓ Howitzer Mobile Secure Enclave Hook License
Dedicated air-gapped deployments, custom hardware enclave integrations, and SLA guarantees.
- ✓ Everything in Pro Tactical
- ✓ Zero-Telemetry Defense-Grade Audits
- ✓ Custom Cloudflare Edge Routing Substrates
- ✓ Multi-Seat Team License Keys & Central Billing
- ✓ Custom Quantization Kernels (NVIDIA / Metal)
- ✓ Priority C++ Engineering Support & SLA
// THE SOVEREIGN FREEMIUM COMMITMENT
Howitzer is built under the philosophy of Aram (அறம்). Core local utility will never be locked behind a paywall. Local GGUF inference, local API debugging, and Bring-Your-Own-Key connections are free forever with zero telemetry.
Pro and Enterprise tiers exist to support high-acuity teams requiring automated dataset synthesis, GPU-accelerated model tuning, multi-model consensus shootouts, and defense-grade key management.
100% In-Memory Isolation. Zero Leaks.
We do not operate servers that log your prompts, store your training datasets, or harvest your API credentials.
Air-Gapped Execution
Local models and API requests run directly from system RAM. Disconnect your Wi-Fi, sever ethernet, and run full inference without a single error.
Hardware Enclave Vault
Frontier API keys are encrypted at rest using AES-256 via macOS Keychain and native Linux secrets daemons. Keys are decrypted strictly in memory upon dispatch.
Audited Test Verification
Verified with LLVM profile coverage across 1,606 lines (100% line coverage) and 8 passing CTest suites. Built to sovereign engineering standards.