Use hosted models without buying a GPU
LU brings local and hosted models into the same desktop app. Use your own hardware for local work, or sign in to LU Cloud for the hosted catalog. Cloud is also available in the browser. Read the differences and limits below before choosing a credit pack or plan.
What cloud can do
- Uncensored text. The catalogue holds 47 chat models. 24 of the 46 cloud chat models we measured answer in full without refusing, and only those carry the No refusals mark. We measured on 2026-09-10, so you do not have to guess. The model that joined after the run carries no mark until it is measured. We add no content filter of our own to text.
- Adult image and video, once you turn your own filter off. Ten image-to-video endpoints and three image models run without a built-in content restriction. Every one of these video endpoints takes an image as its input: generate the still first, then hand it over with one click.
- Chat and coding with the models in the hosted picker. For example, gpt-oss 120B is a Flash chat option, while Qwen3 Coder 480B is a coding-focused option.
- Image and video generation with the modes shown for each model. Editing and input support depend on the selected model, not on a blanket promise that every model supports every tool.
- Model access through the app or an API key. API-key usage draws credits, including on Flash models.
See the current model table and calculator for provider-stated context, quantization where known, image input, tool transport and reasoning controls. A provider's context specification is not a guarantee that every request can use its full native context.
What cloud does not promise
- Not local privacy. Hosted requests leave your machine and use external inference providers. Account or billing location does not establish the inference region.
- Not a free plan. The unmetered Flash path needs an active paid plan. Accounts without an active plan keep paying credits for Flash like for any other model.
- Not unfiltered by default. Media generation follows your account's content policy. The default setting refuses hardcore prompts; you can turn the filter off yourself in the account settings. Text models are not filtered by us at all.
- Never, at any setting. Material involving minors is refused on every request, and you may not upload a photograph of a real, identifiable person without their consent. No setting moves either line.
- Not every local checkpoint. Cloud uses a curated catalog. A model installed on your computer does not automatically become available on our hosted infrastructure.
- Not guaranteed answers or throughput. Models can make mistakes or refuse requests. Request limits, available capacity and credit balances still apply.
Flash chat: a daily allowance, with visible limits
Flash-badged models are unmetered on an active paid plan, inside the apps, up to a published daily ceiling of 500,000 combined input and output tokens per account per day. Only one free request can run at a time. A second simultaneous free request gets a clear busy message, not a silent charge.
What the ceiling means in practice: 500,000 tokens is a long working day of chat, every day, and it resets at 00:00 UTC. It exists so the word unmetered stays true instead of becoming a fair-use paragraph nobody reads. The service reserves a request's token budget before it runs. If the budget cannot fit in the remaining allowance, or the daily allowance has been used, the request follows the normal credit path with a visible notice. Missing final usage can retain the reservation. This is not unlimited API access.
Choose a pack or a plan after checking the calculator
Credit packs start at EUR 15 for 450,000 credits, with no subscription. Pack credits do not expire. Subscription allowances follow the plan's billing rules. Input tokens, reasoning and media settings affect how much work a balance buys.
Check packs and estimated usage
Prefer a plan? Review the Hosted checkout. Review the current terms and privacy policy before purchasing or submitting sensitive information.
Prefer your own hardware?
The desktop app supports local backends on Windows and Linux. There is no Mac build, so on a Mac the hosted studio in the browser is the route. Hardware, model format and backend compatibility determine what runs. Local inference does not make optional cloud features or integrations offline.
Read the local setup guide or check the available desktop downloads.