AI Model Orchestration
Map Open-Source AI Models to physical GPU clusters and deploy routing configs.
Routing Topologies
gemma3:1b Memory Refilter
ollama/gemma3:1b
Load Balanced Endpoints
http://10.171.0.23:11434
qwen2.5:14b
ollama/qwen2.5:14b
Load Balanced Endpoints
http://10.171.0.23:11434
qwen2.5-vl-72b-awq
openai/qwen2.5-vl-72b-awq
Load Balanced Endpoints
http://10.171.0.100:8000
Auto-Discovered Infrastructure Endpoints
Ollama: gemma3:1b
ai-worker-judge 01
http://10.171.0.23:11434
Active in Router Topology
Ollama: qwen2.5:14b
ai-worker-judge 01
http://10.171.0.23:11434
Active in Router Topology
SGLang: qwen2.5-vl-72b-awq
Worker Node 1
http://10.171.0.100:8000
Active in Router Topology
SGLang: qwen2.5-vl-72b-awq
Worker Node 2
http://10.171.0.101:8000
LiteLLM Gateway Sync
Generate the routing configuration from the active models on the left, push it to the remote LiteLLM Router (`10.171.0.202`), and restart the load balancer container.
Deploy Terminal
Awaiting deployment commands...