LU Cloud

The same app, with the heavy work on hosted GPUs. This chapter covers the account, what a plan and a credit pack are, the Flash chat allowance, the marks in the picker, the content policy setting and the API key.

What LU Cloud is

LU Cloud is the same app with the heavy work moved to GPUs that LU Labs rents. Chat runs on a hosted catalogue of open weight models, the Create lanes render on hosted cards, and the two lanes that never run locally (Upscale and Erase Object) run there. You use it when your card is too small, when a model is too big for any desktop, or on a machine with no card at all, and you can mix: local for chat, hosted for video, whichever way round you like.

Local features never need an account. Settings says so under LU Cloud Account: "Sign in to render images, video and chat on LU's cloud GPUs. Local features never need an account."

The account

1 Create the account

Open Settings, General, LU Cloud Account and click "Create account" (or "Sign in" if you have one). Both open your browser on lu-labs.ai. An account is free and needs an email address and a password, or a Google or GitHub login. The desktop app waits with "Waiting for the browser…" and a "Cancel" button; after the login the browser page says "Signed in" and the app picks the session up on its own.

2 Read the account card

Signed in, the card shows "Plan: (name)" or "No active plan", and "Cloud credits (this billing period)" with the numbers. "Manage subscription" or "View plans" opens the account page on lu-labs.ai, where the plan, the payment method, the invoices and the cancellation live; the desktop app never handles payment itself.

Switching to Cloud and back

The switch sits in the header next to the model picker. Clicking "Cloud" once turns it into "Switch to Cloud?" with the note "Click again to move the whole app to Cloud. Answers are then billed to your lu-labs.ai credits." A second click within six seconds moves the whole app: the picker shows hosted models, the Create lanes render hosted, and the composer carries a "Cloud" chip so you can see which side you are on. The Benchmark and Models entries leave the top bar, because they are local tools. One click on the switch brings you back to Local.

Going to Cloud also releases your local models from memory, the same as closing the window to the tray does, so the card is free while you work hosted. They load again on first use after you come back.

Plans, packs and credits

Hosted work is metered in credits. Every request draws from one shared wallet: chat at each model's own rate per token, images and clips at each model's rate per picture or per second. There are two ways to fill the wallet:

What the plans and the packs cost, and how many credits each carries, is on the pricing page, which reads those numbers from the same source as the checkout. This handbook does not repeat them, because they can change and a stale number here would be worse than a link.

The account card in Settings shows what is left this period, and the same figure sits on the account page at lu-labs.ai with a "Buy more credits" link when it runs low. When the wallet is empty a request is refused with a sentence that says so and points at the credits tab of the pricing page; nothing is billed beyond what you loaded.

Flash chat: the daily allowance

Some hosted chat models belong to a Flash class that costs no credits inside the apps on an active paid plan, up to a published daily ceiling of 500,000 input and output tokens per account per day, which resets at 00:00 UTC. The picker marks them with "No credits". The word Flash in a model name does not decide it. The notice in the app reads "Flash: 500,000 input and output tokens per UTC day without credits" and adds "API keys always use credits."

The conditions, as the app and the pricing page state them: it needs an active paid plan, so accounts without an active plan keep paying credits, and a credit pack on its own does not open it; only one free request can run at a time per account, and a second one is refused rather than billed ("One free chat request is already running on this account. Wait for it to finish or stop it first."); a request whose budget would exceed what is left of the day's allowance runs at normal credit pricing, with a notice in the app; and a request through an API key always uses credits, Flash or not. The current list of Flash models is on the pricing page.

The marks in the picker

In Cloud mode the model picker carries two small marks next to some models:

The content policy setting

Settings, General, Content policy has three options, and the text above them reads "Applies to images and video rendered in the cloud. Text is unaffected, and nothing on your own machine is." Signed out, the section reads "Sign in to your LU Cloud account to change this setting."

OptionHint in the app
Strict"The tightest filter we have. Choose this if others use your screen."
Standard"The default for every account."
Off"No filter beyond the legal limits below."

Under all three options one paragraph stays: "Material involving minors, and photographs of real people uploaded without their consent, are refused on every request whatever this is set to, and reported."

So: hosted text is never filtered by this setting, local work is never touched by it, and the two fixed lines apply everywhere.

API keys and the endpoint

Your plan's chat models can be used from other programs. Settings, General, Cloud API Keys reads: "Use the chat models of your plan from Aider, LibreChat or any OpenAI-compatible tool. Base URL https://lu-labs.ai/api/inference/v1, your key as the API key. A key spends plan tokens only; it can never read or change the account." The button "Generate API key on lu-labs.ai" opens the account page, where the section "Cloud API keys" has "Create API key". The key is shown once ("Copy this key now. It will never be shown again."), starts with lu_, and up to five keys can be active; "Revoke this key" ends one.

The endpoint speaks the OpenAI chat API: GET /api/inference/v1/models lists the catalogue, POST /api/inference/v1/chat/completions answers, streaming by default, with the key as a Bearer token. Requests through a key spend credits at the model's rate, Flash or not. The endpoint has a burst limit of 60 requests per minute per account.

The same account in a browser

Everything hosted also runs in a browser at lu-labs.ai/app with the same account, the same credits and the same studios, on any computer including a Mac. That is the way to use it on a machine where the desktop app does not run. The hosted handbook covers that side.

When it goes wrong

The Cloud switch does nothing. It needs two clicks within six seconds. The first turns it into "Switch to Cloud?", the second switches.

"Cloud mode shows hosted models only" in the picker. That is what Cloud mode does. Switch back to Local for your own models.

The request is refused for credits. The wallet is empty for this period: "You're out of credits for this month." followed by where to load up. Buy a pack or wait for the plan to refill on its renewal date; the account card shows both.

"One free chat request is already running on this account. Wait for it to finish or stop it first." Flash allows one free request at a time. Stop the running one (the "Agents · N" panel from chapter 5 lists background agents that may be the one running), or wait.

A Flash model spent credits. One of three reasons: the day's allowance was used up (the app shows a notice), the request came through an API key, or the account has no active paid plan.

A hosted image or video came back refused. Check Settings, General, Content policy. Standard is the default. The two fixed lines apply regardless.

"Sign-in was cancelled in the browser." or the app keeps "Waiting for the browser…". Finish the login in the browser tab the app opened; if that tab is gone, press "Cancel" and "Sign in" again.

A hosted result is gone from the gallery. Hosted results are stored for 7 days and then deleted. Download what you want to keep.


Previous chapter: Create. Next chapter: Settings and troubleshooting. Back to the handbook overview.

Current Hosted offer

Hosted includes 11 Flash models; Pro and Max include 12. Up to 500,000 combined input and output tokens per account per UTC day in app sessions. GLM 5.3 Flash uses credits on Hosted and joins the daily allowance on Pro and Max. API keys always use credits.

EUR 19/month, 900,000 shared monthly credits: up to 900 Neta Lumina Spicy images or 90 five-second LTX 2.3 Spicy clips. Every output spends the same wallet.

Each maximum spends the entire shared monthly wallet on one model. These are alternatives, not separate allowances. Input-image generation, character training, LoRA routes and additional operations cost extra. Spicy video uses an input image and five-second clips. Failed jobs and changing catalog prices can affect results.

Image models and exact quantities · Video models and exact quantities · Venice comparison · Current facts JSON