Use hosted models without buying a GPU

LU brings local and hosted models into the same desktop app. Use your own hardware for local work, or sign in to LU Cloud for the hosted catalog. Cloud is also available in the browser. Read the differences and limits below before choosing a credit pack or plan.

See what it costs

What cloud can do

See the current model table and calculator for provider-stated context, quantization where known, image input, tool transport and reasoning controls. A provider's context specification is not a guarantee that every request can use its full native context.

What cloud does not promise

Flash chat: a daily allowance, with visible limits

Flash-badged models are unmetered on an active paid plan, inside the apps, up to a published daily ceiling of 500,000 combined input and output tokens per account per day. Only one free request can run at a time. A second simultaneous free request gets a clear busy message, not a silent charge.

What the ceiling means in practice: 500,000 tokens is a long working day of chat, every day, and it resets at 00:00 UTC. It exists so the word unmetered stays true instead of becoming a fair-use paragraph nobody reads. The service reserves a request's token budget before it runs. If the budget cannot fit in the remaining allowance, or the daily allowance has been used, the request follows the normal credit path with a visible notice. Missing final usage can retain the reservation. This is not unlimited API access.

Choose a pack or a plan after checking the calculator

Credit packs start at EUR 15 for 450,000 credits, with no subscription. Pack credits do not expire. Subscription allowances follow the plan's billing rules. Input tokens, reasoning and media settings affect how much work a balance buys.

Check packs and estimated usage

Prefer a plan? Review the Hosted checkout. Review the current terms and privacy policy before purchasing or submitting sensitive information.

Prefer your own hardware?

The desktop app supports local backends on Windows and Linux. There is no Mac build, so on a Mac the hosted studio in the browser is the route. Hardware, model format and backend compatibility determine what runs. Local inference does not make optional cloud features or integrations offline.

Read the local setup guide or check the available desktop downloads.

Current Hosted offer

Hosted includes 11 Flash models; Pro and Max include 12. Up to 500,000 combined input and output tokens per account per UTC day in app sessions. GLM 5.3 Flash uses credits on Hosted and joins the daily allowance on Pro and Max. API keys always use credits.

EUR 19/month, 900,000 shared monthly credits: up to 900 Neta Lumina Spicy images or 90 five-second LTX 2.3 Spicy clips. Every output spends the same wallet.

Each maximum spends the entire shared monthly wallet on one model. These are alternatives, not separate allowances. Input-image generation, character training, LoRA routes and additional operations cost extra. Spicy video uses an input image and five-second clips. Failed jobs and changing catalog prices can affect results.

Image models and exact quantities · Video models and exact quantities · Venice comparison · Current facts JSON