Tell it what your post is about, show it the photo, or both. A local AI writes ready-to-edit captions for X, Instagram, LinkedIn, and Facebook. Nothing you type or upload leaves this device.
The caption is somehow always harder than the photo. You know what the post is about; you just cannot make it sound right for the platform, and what works on LinkedIn dies on Instagram. The AI caption tools that fix this send your photo and your half-formed thoughts to their servers, which is a strange trade for something as personal as your camera roll or as commercially sensitive as an unannounced product shot.
This tool does the same job without the upload. Microsoft's Florence-2 vision model looks at your photo in your browser and writes a description you can correct, then a local language model (Qwen 2.5 or Phi 3.5, running on your graphics chip through WebLLM, or on your CPU when WebGPU is unavailable) writes a separate caption for each platform you pick, in the tone and length you choose. The models download once, cache in your browser, and work offline afterward. Your photo, your notes, and your captions never reach a server.
Language models are excellent at rhythm and structure and unreliable about facts. They can misread a photo, invent a detail you never gave them, or land on a tone you would never use with your audience. That is why this tool puts a warning next to every result: read the caption, check anything specific it claims, and edit it into your own voice before you publish. A caption that is 90% written and 100% checked is the honest promise here; no AI, local or cloud, should be posting for you unread.
Type a quick brain dump of what your post is about, add the photo if you have one, pick Instagram (and any other platforms), choose a tone and length, and press the button. A language model running in your own browser writes the caption, with the hook up front and hashtags where Instagram expects them. There is no account, no quota, and no watermark.
No. Both AI models download to your browser once and run on your device: Microsoft's Florence-2 looks at the photo and writes a text description locally, and the caption model writes from that description plus your notes. The photo, the description, and your captions never touch a server, which matters when the photo is of your family, your product, or your clients.
Yes. Select any combination of the four platforms and one click writes a separate caption for each, following each platform's conventions: X stays under 280 characters, Instagram and LinkedIn lead with a hook that survives the see-more fold, LinkedIn keeps hashtags to a professional handful, and Facebook stays conversational with few or no hashtags.
Because language models can be confidently wrong. They can misread a photo, invent details you never mentioned, or strike a tone you would not use. Treat every caption here as a strong first draft: read it, check any facts, names, and numbers, and edit it into your own voice before you publish. The tool reminds you of this next to every result on purpose.
Yes. Drop the photo and the vision model describes what it sees, in an editable box so you can fix anything it gets wrong or add context it cannot know, like names or the occasion. The caption model then writes from that description. You get better results if you add even one line of your own, because the AI cannot know why the moment mattered.
A recent laptop or desktop is the sweet spot: with WebGPU (current Chrome, Edge, Firefox, or Safari) the models run on your graphics chip and captions appear in seconds, and without WebGPU a compact CPU model takes over. Modern phones work too, in a lighter way: phone browsers cap a page's memory well below desktop, so a compact model writes the captions (read them extra carefully) and you type a short photo description yourself instead of the photo AI. The one-time model downloads are cached, so later visits are fast and work offline.