LiteLLM
The lab runs a LiteLLM proxy that gives you access to large language models running on the lab's GPU server (orca), using an OpenAI-compatible API. This lets you use tools like Python scripts, curl, and Claude Code with local open-source models without needing an external API account.
Getting access
Email adarsh@arizona.edu to request an API key. Include a brief description of how you plan to use it.
The API base URL is: https://litellm.lab.pyarelal.xyz
Once you have a key, set it as an environment variable so it persists across sessions. Add this to your shell config file (e.g. ~/.bashrc, ~/.zshrc):
export LITELLM_API_KEY=sk-...
Then reload your shell: source ~/.bashrc (or open a new terminal).
Available models
To see which models are currently available:
curl https://litellm.lab.pyarelal.xyz/models \
-H "Authorization: Bearer $LITELLM_API_KEY"
Models are named <family>:<size>[-a<N>b]-<quant>, e.g. qwen3.6:35b-a3b-q8_0. The name tells you three things: total parameter count (35b), whether it's a mixture-of-experts model (a3b = only 3B parameters active per token; no a-suffix means dense), and the quantization level (q8_0 = 8-bit, near-lossless; q4_k_m = 4-bit). Dense models are generally stronger per total parameter; MoE models generate faster for their size.
If you are unsure which model to use, use qwen3.8:27b-q8_0 — it is the strongest general model we serve. It serves 4 requests concurrently (each up to the full 262k context window, drawn from a shared pool), and repeated requests with a shared prefix — an ongoing chat, a system prompt — skip reprocessing what the server has already seen, even if other requests ran in between. It stays loaded while in use and for 1 idle hour afterwards, so responses are normally instant; the first request after a longer idle gap takes about a minute while the model reloads. All other models share the remaining GPU memory one at a time, so requesting one that is not currently loaded pays a model-load delay (up to ~2 minutes for the largest), and each unloads after 1 idle hour.
Using with curl
curl -X POST https://litellm.lab.pyarelal.xyz/chat/completions \
-H "Authorization: Bearer $LITELLM_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "qwen3.8:27b-q8_0",
"messages": [{"role": "user", "content": "Hello!"}]
}'
Using with Python
Install the OpenAI SDK if you don't have it: pip install openai
import os
from openai import OpenAI
client = OpenAI(
api_key=os.environ["LITELLM_API_KEY"],
base_url="https://litellm.lab.pyarelal.xyz",
)
response = client.chat.completions.create(
model="qwen3.8:27b-q8_0",
messages=[{"role": "user", "content": "Hello!"}],
)
print(response.choices[0].message.content)
Using with Claude Code
You can use Claude Code with the lab's models by pointing it at LiteLLM instead of Anthropic's API. Set these environment variables before running claude:
export ANTHROPIC_API_KEY=$LITELLM_API_KEY
export ANTHROPIC_BASE_URL=https://litellm.lab.pyarelal.xyz
claude
Then switch to a lab model inside Claude Code with the /model command:
/model qwen3.8:27b-q8_0
Note: open-source models have different capabilities than Claude — some Claude Code features (e.g. complex tool use) may not work as well.
Using with omp (Oh My Pi)
omp is an open-source coding-agent CLI that works well with local models. Point it at the lab proxy by adding a provider block to ~/.omp/agent/models.yml:
providers:
litellm:
baseUrl: https://litellm.lab.pyarelal.xyz/v1
apiKey: LITELLM_API_KEY
api: openai-completions
authHeader: true
models:
- id: qwen3.8:27b-q8_0
name: Qwen 3.8 27B Q8
contextWindow: 262144
maxTokens: 32768
reasoning: true
input: ["text", "image"]
Details that matter:
apiKeynames the environment variable holding your key (theLITELLM_API_KEYyou set above) — don't paste the key itself into the file.idmust exactly match a model id from the/modelslisting; add one entry per model you want to use.contextWindowmust equal the server-side context for that model — 262144 for everything currently served exceptgemma4:e4b-q8_0(32768). If you declare a smaller value, omp starts compacting (summarizing away) your conversation far earlier than necessary.input: ["text", "image"]belongs only on the vision-capable models (all of the current lineup exceptqwen3-coder-next).
Inside omp, select the model with /model and pick litellm/qwen3.8:27b-q8_0. If you assign omp's model roles (default/smol/plan/…) explicitly, use at most one model other than qwen3.8:27b-q8_0 and gemma4:e4b-q8_0 across all roles: those two stay loaded independently, but every other model loads one at a time on the server, so a session whose roles alternate between two of those pays a model-swap delay of a minute or more on every switch.