Why Pay for Big Tech When You Can Have Choice?
The AI market is changing fast. Open-source models now rival the best proprietary systems at a fraction of the cost. General Bots puts you at the forefront of this shift: self-hosted, private, open source, and completely free from vendor lock-in.
GLM-5.2 — Zhipu AI
Zhipu AI's latest flagship — 745B MoE with 1M-token context, trained entirely on Huawei Ascend chips. Scores 94.8% on GSM8K, 88.7% MMLU, and ranks among the top open-source models on SWE-Bench Pro. Fully open weights, MIT license, self-hostable.
Qwen 3.7 — Alibaba Cloud
Alibaba's strongest model family — Qwen3.7-Max scores 56.6 on Intelligence Index v4.0, #1 among Chinese models. 1M-token context, 50.8% Terminal-Bench Hard, lowest hallucination rate at 22.9%. Available via API and open weights under Apache 2.0.
DeepSeek V4-Pro
The next-generation open-weight MoE flagship — 1.6T total params, 49B active, 1M-token context. Scores 80.6% on SWE-Bench Verified (highest open-weights entry). MIT license, self-hostable. V4-Flash variant offers 284B params at $0.14/M input tokens for cost-sensitive workloads.
Kimi K2.7 Code — Moonshot AI
Moonshot AI's latest coding-focused 1T-MoE model. Ties GPT-5.5 on SWE-Bench Pro at 5-6x lower cost. Agent Swarm primitive spawns up to 300 sub-agents across 4,000 coordinated steps. 262K-token context with auto-compression. Open weights under MIT.
MiniMax M3
Shanghai-based MiniMax's frontier open-weight model — first to combine 1M-token context, native multimodality (image + video), and 59.0% SWE-Bench Pro (surpassing GPT-5.5). MiniMax Sparse Attention delivers 9x faster prefill. Open weights available.
Plus: Yi, Step, Doubao, ERNIE
The Chinese AI ecosystem is vast. Yi (01.AI) delivers competitive coding performance, Step (StepFun) excels at multimodal reasoning, Doubao (ByteDance) leads in consumer AI, and ERNIE (Baidu) offers deep enterprise integration. Add any OpenAI-compatible API — one interface, every model, your rules.
Cost Comparison
Open-source models deliver world-class performance at near-zero marginal cost. No per-seat fees, no per-token pricing, no data training on your prompts.
| Model | Cost per 1M tokens (input) | Self-Hostable | Data Privacy |
|---|---|---|---|
| DeepSeek V4-Pro | $0.44 | Yes | Yes |
| DeepSeek V4-Flash | $0.14 | Yes | Yes |
| Qwen 3.7-Max | $2.50 | No (API) | Yes (API) |
| GLM-5.2 | $0.11 | Yes | Yes |
| Kimi K2.7 | $0.95 | No (API) | Yes (API) |
| MiniMax M3 | $0.30 | Yes | Yes |
High-Performance Orchestration
General Bots' LLM orchestrator is built in Rust, not Python. This means zero-copy token handling, lock-free concurrent request processing, and memory-safe execution. The orchestrator handles thousands of simultaneous streaming requests across multiple model providers with sub-millisecond routing overhead.
- Dynamic Model Routing — Route simple queries to cheap models, complex ones to frontier models
- Intelligent Failover — If your primary provider goes down, switch to a backup in milliseconds
- Semantic Caching — Cache semantically similar queries, reducing costs by up to 90%
One API. Every Model. Your Choice.
DeepSeek, Qwen, GLM, Kimi, MiniMax, Yi — run the best open models on your own infrastructure. No vendor lock-in. No per-seat fees. Complete data privacy.