5.3 KiB
AI setup — free & local (no paid Gemini required)
The gallery's AI splits into two groups. Group 1 already runs for free, on-device, with no API key. Only Group 2 (generative pixel editing) ever needed Gemini — and that backend is now swappable to a free or fully-local model.
Group 1 — Analysis & background removal — FREE, LOCAL, no key ✅
These run entirely in the user's browser (models download once, then cached). Nothing is sent to any server; no key or config is required. This is the bulk of "AI" in the app:
| Capability | Runs on | Library |
|---|---|---|
| Object detection → tags, Objects browser, auto object albums | in-browser (WebGL/WASM) | TensorFlow.js COCO-SSD |
| Face detection + recognition → People clustering | in-browser | face-api.js |
| OCR (text in images) → search + Documents album | in-browser | tesseract.js |
| Semantic search / captions embeddings | in-browser | transformers.js (CLIP) |
| Remove Background (editor) | in-browser (WASM) | @imgly/background-removal |
So imported photos get the same analysis as captured ones, and Remove Background works, without Gemini or any key. If you deploy with no AI keys at all, everything here still works.
Group 2 — Generative pixel editing — pick a free or local backend
These edits generate new pixels from a prompt: Restore, Colorize, Replace Sky, and free-form
Prompt. Pick a backend with the AI_EDIT_PROVIDER env var (default auto). The server route
(/api/ai/edit) proxies to it so no key ever reaches the browser.
auto chooses the first configured of: local → huggingface → gemini.
Option 1 — Your own local Stable Diffusion (free + private, recommended) 🏆
Run Stable Diffusion locally with any of AUTOMATIC1111 WebUI, Forge, or SD.Next — they all
expose the same img2img HTTP API. Start it with the API enabled:
# AUTOMATIC1111 example
./webui.sh --api --listen # serves http://127.0.0.1:7860
Then set on the app server:
AI_EDIT_PROVIDER=local
LOCAL_SD_URL=http://127.0.0.1:7860
# optional tuning:
LOCAL_SD_DENOISE=0.55 # 0=keep original … 1=fully reimagine
LOCAL_SD_STEPS=25
LOCAL_SD_SAMPLER=Euler a
100% free, unlimited, private (nothing leaves your machine), and works for every generative op. Needs a machine with a GPU (or a slow CPU). This is the best "own model, everything free" path.
Deploying to Vercel? A serverless function can't reach your
localhost. Either run the whole app on the same box as SD (Option B in DEPLOY-MANUAL.md), or expose your SD box over a tunnel (e.g.cloudflared,ngrok) and pointLOCAL_SD_URLat the public URL.
Option 2 — Hugging Face Inference API (free hosted, zero GPU)
Free token, no GPU needed. Uses an instruction image-editing model.
- Create a free token at https://huggingface.co/settings/tokens (read scope is fine).
- Set:
AI_EDIT_PROVIDER=huggingface HF_API_TOKEN=hf_xxxxxxxx HF_IMAGE_MODEL=timbrooks/instruct-pix2pix # optional; instruction-based image edit
Caveats of the free tier: the first call may return 503 "model loading" (retry in ~20s), and
there are rate limits. Great for a demo; for heavy use prefer Option 1. If a model becomes gated, pick
another image-to-image model with HF_IMAGE_MODEL.
Option 3 — Google Gemini (needs a billed key for images)
The free Gemini tier only returns text — image output (gemini-2.5-flash-image / "Nano Banana")
requires billing enabled, which is why it fails for you. If you enable billing:
AI_EDIT_PROVIDER=gemini
GEMINI_API_KEY=...
GEMINI_IMAGE_MODEL=gemini-2.5-flash-image # optional
Option 4 — none
Leave all three unset (or AI_EDIT_PROVIDER=none). Generative edits show a friendly "not configured"
message; all of Group 1 (analysis + Remove Background) still works, plus every non-AI editor tool
(crop, rotate, filters, adjust, annotations, and the whole video editor).
Bring your own provider (SDK-level)
The SDK takes a pluggable AIProvider (ai prop on <PhotoGallery>). Every method is optional:
interface AIProvider {
detectObjects?(item, image): DetectedObject[] | Promise<…>;
detectFaces?(item, image): DetectedFace[] | Promise<…>;
ocr?(item, image): string | Promise<string>;
embedImage?(item, image): number[] | Promise<number[]>;
embedText?(text): number[] | Promise<number[]>;
generativeEdit?(item, image, op): Blob | Promise<Blob>; // return the edited image
}
The demo's provider (apps/web/src/lib/ai/createDemoAIProvider.ts) wires the in-browser models for
analysis and background-removal, and routes the other generative ops to /api/ai/edit. To use a
different service (local Ollama+vision, ComfyUI, Replicate, your own model server, etc.), implement
generativeEdit (and/or the analysis methods) and pass your provider — no changes elsewhere.
Summary
| Task | Free? | How |
|---|---|---|
| Object/face/OCR/embedding analysis | ✅ always free | in-browser, no key |
| Remove Background | ✅ always free | in-browser (@imgly) |
| Filters / crop / rotate / adjust / annotate / video editor | ✅ always free | no model at all |
| Restore / Colorize / Replace-Sky / Prompt edit | ✅ free via local SD or HF free token | AI_EDIT_PROVIDER |