# LiteLLM

The lab runs a [LiteLLM](https://litellm.ai) proxy that gives you access to large language models running on the lab's GPU server (orca), using an OpenAI-compatible API. This lets you use tools like Python scripts, curl, and Claude Code with local open-source models without needing an external API account.

## Getting access

Email <adarsh@arizona.edu> to request an API key. Include a brief description of how you plan to use it.

The API base URL is: `https://litellm.lab.pyarelal.xyz`

Once you have a key, set it as an environment variable so it persists across sessions. Add this to your shell config file (e.g. `~/.bashrc`, `~/.zshrc`):

```bash
export LITELLM_API_KEY=sk-...

```

Then reload your shell: `source ~/.bashrc` (or open a new terminal).

## Available models

To see which models are currently available:

```bash
curl https://litellm.lab.pyarelal.xyz/models \
  -H "Authorization: Bearer $LITELLM_API_KEY"

```

Models are named `<family>:<size>[-a<N>b]-<quant>`, e.g. `qwen3.6:35b-a3b-q8_0`. The name tells you three things: total parameter count (`35b`), whether it's a mixture-of-experts model (`a3b` = only 3B parameters active per token; no `a`-suffix means dense), and the quantization level (`q8_0` = 8-bit, near-lossless; `q4_k_m` = 4-bit). Dense models are generally stronger per total parameter; MoE models generate faster for their size.

If you are unsure which model to use, use **qwen3.8:27b-q8\_0** — it is the strongest general model we serve. It serves 4 requests concurrently (each up to the full 262k context window, drawn from a shared pool), and repeated requests with a shared prefix — an ongoing chat, a system prompt — skip reprocessing what the server has already seen, even if other requests ran in between. It stays loaded while in use and for 1 idle hour afterwards, so responses are normally instant; the first request after a longer idle gap takes about a minute while the model reloads. All other models share the remaining GPU memory one at a time, so requesting one that is not currently loaded pays a model-load delay (up to ~2 minutes for the largest), and each unloads after 1 idle hour.

## Using with curl

```bash
curl -X POST https://litellm.lab.pyarelal.xyz/chat/completions \
  -H "Authorization: Bearer $LITELLM_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "qwen3.8:27b-q8_0",
    "messages": [{"role": "user", "content": "Hello!"}]
  }'

```

## Using with Python

Install the OpenAI SDK if you don't have it: `pip install openai`

```python
import os
from openai import OpenAI

client = OpenAI(
    api_key=os.environ["LITELLM_API_KEY"],
    base_url="https://litellm.lab.pyarelal.xyz",
)

response = client.chat.completions.create(
    model="qwen3.8:27b-q8_0",
    messages=[{"role": "user", "content": "Hello!"}],
)
print(response.choices[0].message.content)

```

## Using with Claude Code

You can use Claude Code with the lab's models by pointing it at LiteLLM instead of Anthropic's API. Set these environment variables before running `claude`:

```bash
export ANTHROPIC_API_KEY=$LITELLM_API_KEY
export ANTHROPIC_BASE_URL=https://litellm.lab.pyarelal.xyz
claude

```

Then switch to a lab model inside Claude Code with the `/model` command:

```
/model qwen3.8:27b-q8_0

```

Note: open-source models have different capabilities than Claude — some Claude Code features (e.g. complex tool use) may not work as well.

## Using with omp (Oh My Pi)

[omp](https://omp.sh) is an open-source coding-agent CLI that works well with local models. Point it at the lab proxy by adding a provider block to `~/.omp/agent/models.yml`:

```yaml
providers:
  litellm:
    baseUrl: https://litellm.lab.pyarelal.xyz/v1
    apiKey: LITELLM_API_KEY
    api: openai-completions
    authHeader: true
    models:
      - id: qwen3.8:27b-q8_0
        name: Qwen 3.8 27B Q8
        contextWindow: 262144
        maxTokens: 32768
        reasoning: true
        input: ["text", "image"]

```

Details that matter:

- `apiKey` names the *environment variable* holding your key (the `LITELLM_API_KEY` you set above) — don't paste the key itself into the file.
- `id` must exactly match a model id from the `/models` listing; add one entry per model you want to use.
- `contextWindow` must equal the server-side context for that model — **262144** for everything currently served except `gemma4:e4b-q8_0` (32768). If you declare a smaller value, omp starts compacting (summarizing away) your conversation far earlier than necessary.
- `input: ["text", "image"]` belongs only on the vision-capable models (all of the current lineup except `qwen3-coder-next`).

Inside omp, select the model with `/model` and pick `litellm/qwen3.8:27b-q8_0`. If you assign omp's model roles (default/smol/plan/…) explicitly, use at most **one** model other than `qwen3.8:27b-q8_0` and `gemma4:e4b-q8_0` across all roles: those two stay loaded independently, but every other model loads one at a time on the server, so a session whose roles alternate between two of those pays a model-swap delay of a minute or more on every switch.

## Using for speech-to-text

The proxy also serves WhisperX transcription as the model `whisperx:large-v3`, so you can transcribe audio with the same `LITELLM_API_KEY` instead of a separate WhisperX key:

```bash
curl https://litellm.lab.pyarelal.xyz/v1/audio/transcriptions \
  -H "Authorization: Bearer $LITELLM_API_KEY" \
  -F "file=@/path/to/sample.wav" \
  -F "model=whisperx:large-v3"

```

Speaker diarization is on by default; the returned `text` field is prefixed with `[SPEAKER_00]`, `[SPEAKER_01]`, etc. for multi-speaker audio.

The proxy drops non-standard form fields before forwarding, so `response_format=text` and the diarization hints `num_speakers` / `min_speakers` / `max_speakers` have no effect here — you always get the default JSON response. If you need those, use the WhisperX service directly.