TL;DR: nb2lite-skill-claude wraps Google's
gemini-3.1-flash-lite-imagemodel in a tiny FastMCP server and packages it as a Claude Code skill. You type "generate an image of a cyberpunk kitchen" into Claude Code, and it just... does it. Then you say "add a neon RAMEN sign" and it edits the same image without re-prompting the whole scene. Oh, and the cover image of this article? Generated by the thing the article is about — dogfooding all the way down. More on that at the end.
Background: why another image tool?
Most image-generation workflows are stateless. You send a prompt, you get pixels back, and the model immediately forgets everything. Want to tweak the result? You re-describe the entire scene and pray the character, lighting, and composition survive the round trip. (Narrator: they don't.)
Google's Nano Banana 2 Lite — the friendly nickname for gemini-3.1-flash-lite-image — takes a different approach. It's a high-efficiency image model with sub-2-second generations, solid text rendering in 25+ languages, and — the headline feature — support for the stateful Interactions API, which lets you iterate on an image across multiple turns while the model keeps the visual context server-side.
This repo glues that capability into Claude Code, so your coding agent can generate and iteratively refine images as a natural part of a session. It ships as two things in one repo:
- A Model Context Protocol (MCP) server (
nb2lite-agent, a single-file FastMCP app inserver.py) exposing exactly four tools. - A Claude Code skill (
nb2lite-image) that teaches Claude when and how to use those tools well.
The Interactions API: images with a memory
The Interactions API is Gemini's stateful endpoint. The core loop looks like this:
- You call
client.interactions.create(...)with a prompt andstore=True. - The response includes an
interaction_id— a handle to the turn's visual context, persisted on Google's servers. - On the next call, you pass
previous_interaction_id, and the model edits the existing canvas — preserving character, style, lighting, and pixel continuity.
So instead of this (stateless suffering):
"A watercolor fox in a forest at dawn, mist, soft light, wearing a red scarf, three birch trees on the left, and now also holding a lantern"
...you write this:
"Add a lantern in its paw."
That's it. The stored context holds the rest.
A few practical details the server handles for you:
- Every turn returns a new interaction ID. Chain the latest one; editing from a stale ID silently forks your session from an older state (a subtle and very annoying bug if you roll this by hand).
-
Aspect ratio is chosen at generation time (
1:1,16:9,9:16,4:3,3:4) and inherited on stateful edits — changing it mid-session degrades pixel continuity, so the edit tool deliberately doesn't accept one. -
Thinking levels:
low(default, fast drafts) orhigh(complex rendering, accurate text layout, character composition). The generic API spec also listsminimalandmedium, but the live API rejects them for this model with an HTTP 400 — the server saves you from discovering that the hard way.
What is MCP, in one minute
The Model Context Protocol is an open standard for connecting AI assistants to tools and data. Before it, giving a model access to some service meant writing a bespoke integration for each assistant — N assistants × M services, everyone reinventing the same plumbing. MCP collapses that: a tool author writes one MCP server that exposes typed tools, and any MCP-capable client (Claude Code, Claude Desktop, and a growing list of others) can discover and call them with no per-client glue code.
An MCP server is usually a small local process that speaks JSON-RPC over stdio. The client launches it, asks "what tools do you have?", and from then on the model can call them like functions.
The nb2lite-agent server exposes exactly four:
| Tool | What it does |
|---|---|
generate_image |
Text → 1k image. Saves locally, returns the path + an interaction ID. |
edit_image |
Stateful edit: takes the previous interaction ID + a description of only the change. |
edit_local_image |
Uploads any local image file inline (base64) and applies an edit — your entry point for existing files. |
get_help |
Reports live config: API key status, active model, output directory, full tool reference. |
Images land on disk as gen_<timestamp>_<uuid8>.jpg (or edit_/edit_local_ prefixed) — the UUID suffix keeps concurrent generations from clobbering each other. Errors come back as 🔴 ... text strings rather than protocol errors, so the agent can read and react to them.
And what's a Claude Code skill?
If MCP is the hands (the tools Claude can physically call), a skill is the muscle memory — a markdown file (SKILL.md) plus bundled resources that load into Claude's context and teach it the workflow: which tool to reach for, in what order, with which constraints.
For nb2lite-image, the skill encodes things like:
- Call
get_helpfirst when diagnosing setup issues — if the API key is missing, nothing else will work. - Keep edit prompts incremental: describe the change, not the scene.
- Always chain the latest interaction ID.
- Generations are billable — batch related edits and prefer
thinking_level: lowfor drafts.
The skill also bundles the MCP server itself (mcp/server.py), its requirements, an installer script, and a vendored copy of the Interactions API developer guide — so it's self-contained: install the skill, and you have everything needed to also stand up the server.
Installing it: the "I just want it to work" edition
You need three things: Python 3.10+, Claude Code, and a Gemini API key (free from Google AI Studio). Pick one of the paths below.
Path A: The plugin marketplace (fewest keystrokes)
Inside Claude Code, type:
/plugin marketplace add xbill9/nb2lite-skill-claude
/plugin install nb2lite-image@nb2lite-skill-claude
This installs the skill and auto-registers the MCP server. The plugin manifest carries no API key (as it should!) — the server reads GEMINI_API_KEY from your environment, so make sure it's exported before launching Claude Code.
Path B: Clone and bootstrap (this repo)
# 1. Get the code
git clone https://github.com/xbill9/nb2lite-skill-claude.git
cd nb2lite-skill-claude
# 2. One-command setup: installs deps, registers the MCP server
# in .mcp.json, and prompts for your API key (stored in ~/gemini.key)
./init.sh
# 3. Restart Claude Code in this directory and approve the server
# when prompted. Verify with:
/mcp # should list nb2lite-agent
That's genuinely it. init.sh is safe to rerun if anything looks off.
Path C: Install into your project
From a clone of the repo:
make init TARGET=/path/to/your/project ARGS='--output-dir ./images'
This copies the skill into <project>/.claude/skills/nb2lite-image/ and writes the nb2lite-agent entry into that project's .mcp.json. It reuses ~/gemini.key if you've set one up. Restart Claude Code in the target project, approve the server, done.
Path D: Docker (nothing on the host but Docker)
The server is published as xbill9/nb2lite-agent:
claude mcp add nb2lite-agent --env GEMINI_API_KEY="$(cat ~/gemini.key)" -- \
docker run --rm -i -e GEMINI_API_KEY -v "$PWD:$PWD" -w "$PWD" xbill9/nb2lite-agent
The -v "$PWD:$PWD" -w "$PWD" mount matters: the server saves images to disk and reads local files for edit_local_image, so the container must see your project at the same absolute path as the host.
Troubleshooting, the whole guide
-
/mcpdoesn't list the server → restart Claude Code in the project directory. - Tools return
🔴 GEMINI_API_KEY is not set→ runsource set_env.sh(or export the key) and restart. - Anything else → ask Claude to call
get_help; it reports the live config.
Examples: a session in practice
Once installed, you talk to it in plain English. A real flow looks like:
You: "Generate a cozy cabin in a snowy forest at dusk, 16:9."
Claude calls:
generate_image(
prompt="A cozy log cabin in a snowy forest at dusk, warm light in the windows",
aspect_ratio="16:9",
thinking_level="low",
)
# 🟢 Saved to: ./gen_1784759001_a1b2c3d4.jpg
# Interaction ID: v1_ChdpRU5...
You: "Nice. Add smoke curling from the chimney."
edit_image(
previous_interaction_id="v1_ChdpRU5...",
edit_prompt="add gentle smoke curling from the chimney",
)
# 🟢 Saved to: ./edit_1784759050_e5f6a7b8.jpg
# Interaction ID: v1_Xk9mPq2... ← a NEW id; the next edit chains this one
You: "Now make it night, with aurora in the sky."
Same tool, newest ID, and the cabin, trees, and chimney smoke all stay put — only the sky changes. No re-prompting, no continuity roulette.
And for images that didn't come from the model at all:
You: "Take ./whiteboard-sketch.png and render it as a clean 3D product mockup."
edit_local_image(
image_path="./whiteboard-sketch.png",
edit_prompt="render this hand-drawn sketch as a high-fidelity 3D product mockup",
aspect_ratio="4:3",
)
It returns an interaction ID too — so follow-up refinements switch to edit_image and go stateful from there.
Dogfooding: about that cover image 🐕🍖
If the term is new to you: "eating your own dog food" means using your own product for real work, not just demoing it. It's the difference between "this should work" and "I ship with this every day." If a tool is good enough for your users, it should be good enough for you — and if it isn't, you'll be the first to feel the pain and fix it.
This repo dogfoods itself at every layer:
- The skill is active inside its own repository — open Claude Code in a clone and the
nb2lite-imageskill andnb2lite-agentserver are already wired up, so every development session doubles as an integration test. - The integration tests (
make test) drive the same four MCP tools an end user would, against the live API. - And now, the cover image of this article was generated by the exact skill the article describes, from inside a Claude Code session in this repo. One tool call, first attempt, no retouching:
generate_image(
prompt="A wide tech blog cover illustration: a friendly robot artist "
"painting a glowing galaxy on an easel, while a chain of connected "
"frames behind it shows the same picture evolving step by step "
"(day sky, then sunset, then storm with lightning). Flat vector "
"style, deep indigo background, neon cyan and orange accents. "
"Title text 'NB2Lite + MCP', subtitle 'Stateful image editing "
"as a Claude Code skill'. Crisp, accurate lettering.",
aspect_ratio="16:9",
thinking_level="high",
)
# 🟢 Image successfully saved!
# • Saved to: gen_1784759177_cbab8b65.jpg
# • Interaction ID: v1_ChdpRU5hb2o3SWMzV2pNY1AtUFgy...
(That exact output is committed to the repo as devto-cover.jpg, receipts and all.)
Worth noticing:
-
The text rendered correctly. "NB2Lite + MCP" and the full subtitle came out crisp and typo-free — that's what
thinking_level: "high"buys you on text-heavy layouts. - The model illustrated its own pitch. The chain of frames (day → sunset → storm → galaxy) is the stateful edit loop — the image explains the Interactions API better than a diagram I'd have drawn by hand.
-
If I wanted the accent color changed, I wouldn't regenerate — I'd
edit_imagewith that interaction ID and say "make the orange accents magenta." That's the whole point.
Dogfooding is the cheapest credibility there is: no cherry-picked gallery, no "results may vary" fine print — the tool's real output is literally the first thing you saw when you opened this article. If the skill had flubbed the lettering or mangled the layout, you'd be looking at the evidence right now. Instead, the article ships with its own proof baked into the header.
Links
- Repo: github.com/xbill9/nb2lite-skill-claude (Apache-2.0)
- Docker image: hub.docker.com/r/xbill9/nb2lite-agent
- Interactions API reference: ai.google.dev/api/interactions-api
- Model Context Protocol: modelcontextprotocol.io
This is a third-party community project, not affiliated with or endorsed by Anthropic or Google. Bring your own Gemini API key — and remember generations are billable, so draft on low and save high for the money shot.

Top comments (0)