Reference
Models and credits
The model list and how Modloom Credits are spent.
You choose a model for each request, from fast and cheap to the heaviest available.
The models#
| Model | Provider | Free plan |
|---|---|---|
| GPT-6.1 Sol | OpenAI | Paid plans |
| Claude Opus 5.5 | Anthropic | Paid plans |
| Claude Sonnet 5.5 | Anthropic | Paid plans |
| Claude Sonnet 5 | Anthropic | Paid plans |
| Claude Haiku 4.5 | Anthropic | Paid plans |
| GPT-6 Luna | OpenAI | Included |
| GPT-5.6 Sol | OpenAI | Paid plans |
| GPT-5.6 Terra | OpenAI | Paid plans |
| Gemini 3.1 Pro | Paid plans | |
| Gemini 3.8 Flash | Paid plans | |
| Gemini 3.5 Flash Lite | Paid plans | |
| DeepSeek V4.1 Flash | DeepSeek | Included |
| DeepSeek V4 Pro | DeepSeek | Paid plans |
| GLM 5.3 Prime | Z.ai | Paid plans |
| GLM 5.3 | Z.ai | Paid plans |
| GLM 5.3 Flash | Z.ai | Included |
Models marked Paid plans need a Maker or Studio plan. On the free plan they show a padlock in the picker.
Modloom Credits#
AI work is metered in Modloom Credits. There is no unmetered model. One credit is a fixed amount of work on the reference model, and every other model is priced relative to it.
- A model that costs twice as much per token spends twice the credits for the same work.
- The picker shows each model's rate next to its name, for example x1 or x0.05. x1 is the standard rate.
- Both input and output count, and output is weighted more heavily because it costs more.
- Rates follow the provider's current prices and are refreshed daily, so a rate can change over time.
What affects the cost of a run#
- The model you pick.
- How much the agent reads and writes. Large projects and long threads mean more input.
- How many attempts it takes. Builds that fail and get fixed cost more than builds that pass first time.
A few habits keep runs cheap:
- start with a fast model and move up only if the result is not good enough,
- plan first so the build run is shorter,
- keep requests focused on one change.
Limits per run#
Every run has a ceiling on turns, time and cost so a single request cannot run away. If a run reaches it, the run stops and tells you. Split the request into smaller ones.
Long conversations#
When a thread grows past what the model can hold, older parts are summarized automatically. You see a "Conversation compacted" notice. Details from early in a very long thread can be lost, so restate anything important.
When credits run out#
When your weekly credits are used up, runs are blocked until the next reset, and you get an email. See Plans and limits.