Here’s exactly what I typed:
”A cartoon tiger, with a big angular jawline, lounging gracefully with his back legs outstretched, eyeing off a cartoon mouse in the far foreground. It’s golden hour, a rainforest floor and canopy surround them, and a thin boot-carved dirt pathway meanders into the distance where mousey eats cheese, magical and whimsical tone, vibrant and visually striking, outdoors, natural setting, golden hour lighting, warm glow, high quality, detailed.”
What came back was gorgeous. Exactly the tiger. Exactly the light. Exactly the small, comic tension between a very large cat and a mouse who hasn’t noticed him yet.
I sat there pleased with myself for about four seconds. Then a different thought landed, and it was less flattering: nobody else would have written that.
The gap I already knew was there
By this point I’d been fighting Gemini’s early image results for a while - reliably, frustratingly wrong, good enough to see what was possible and nowhere near good enough to use. I’d already found my own way through that, prompt by prompt, the slow way, until sentences like the one above started coming out of me without much effort.
That’s the part that stopped me. It stopped being effort. I could produce forty words of jaw-angle, canopy detail, and lighting terminology practically without thinking, because I’d done it enough times that the vocabulary had become mine.
Most people don’t have that. Not because they’re less capable — because nobody hands you that vocabulary before you need it. You know the image in your head with total clarity. You have no idea which words make an AI see it the way you do.
Recent research behind that gap says something specific: it’s not that people can’t tell a good prompt from a bad one, or can’t describe what they want in plain language. They can do both. What they consistently lack is the style-specific vocabulary — the “golden hour,” the “shallow depth of field,” the particular grammar of image prompting — that turns a clear idea into a workable instruction. It’s a learnable skill. It’s just not one that anyone is taught.
When you know what you want to see, but the AI doesn’t understand your words.
Let’s say that you want to create an image. You have it perfectly clear in your head — the composition, the mood, the feeling. You know exactly what you want to see.
But you don’t know how to tell Claude or Gemini what you want.
The gap between the image in your head and the words you need to make the AI create it — that gap is massive. It’s not about intelligence. It’s about translation. The AI needs specificity. It needs detail. It needs the language of image generation — the prompting vocabulary that actually works.
Early on in my dabbling with AI, I’d write something simple — “a woman at a desk” — and get back something that was technically correct but completely missed the feeling I was after.
I was failing to speak the language.
Then I thought: Claude can help me learn this language. Claude can take my fuzzy, emotional, imprecise idea and translate it into the kind of prompt that actually works with image generation tools.
So I built Mahna Mahna - a simple little desktop app - from a conversation with Claude about my frustration with convincing Gemini to make me the image that I wanted. We chatted about how to embellish my words so the image had more depth or accuracy, how an app should behave when the user starts it up, how the app should output the prompt and what features I’d like to see. No more than 20 minutes of a typed conversation. Et Voila, Claude created me an App.
You tell Mahna Mahna what you want. You describe it in plain language - messy, emotional, imprecise. And it translates that into a prompt that Gemini or any other image generator can actually use. Specific. Detailed. Written in the vocabulary that these tools understand.
Suddenly the gap closes. Suddenly what’s in your head can become what appears on the screen.
That’s when I realized: so much of what’s hard about using AI isn’t about the AI being stupid. It’s about us not speaking its language. And once you learn to translate - once you have a tool that helps you translate - everything changes.
Mahna Mahna is that translator.
Where the name came from
I couldn’t have told you, if you’d asked me directly, why I called it that. It just arrived already attached to the idea, the way names sometimes do.
What I could tell you was the feeling: the Muppets, a stage, curtains opening on something grand and a little ridiculous, the particular showmanship of a full production number promising you’re about to see something worth watching. Vaudeville dressed up as children’s television.
The real history turns out to be stranger than the memory. The tune wasn’t written for the Muppets at all — Italian composer Piero Umiliani wrote it in 1968 for a film about Scandinavian life, and it only found its way to Jim Henson’s puppets a year later, first on Sesame Street, then as the very first thing audiences saw when The Muppet Show premiered in 1976: curtains rising on a scrappy, orange-haired performer trying, and failing, to bring his two reluctant backup singers into line.
That’s the version I actually remember. Not the Italian film. The curtain rising on something that shouldn’t work and somehow does — spectacle, a bit of chaos, the promise that what’s about to happen is worth watching.
Which, now that I’ve written it out, might be the most honest description of what I actually wanted from an AI-generated image the whole time.

What it means for someone without forty words to spare
Here’s the part that actually built the app, though, not just named the problem.
If you’re on a free tier of ChatGPT or Gemini or whatever tool you might try to use to create or edit an image — and a lot of the people this blog is for are — every failed attempt costs you something. You don’t get unlimited tries to discover, the hard way, that “a woman at a desk” gets you something flat and generic while “a woman in her forties, warm afternoon light, mid-thought, gaze turned toward the window” gets you something that actually feels like a moment. You might get three or four goes before the day’s limit is gone, and you’ve spent all of them relearning a vocabulary I’d already built by accident.
So Mah-nah Mah-nah doesn’t generate the image. It sits in front of the image generator and does the translation — takes the fuzzy, honest, everyday version of what you want and hands back the specific, detailed version an image tool actually needs. You still don’t have to know what “golden hour” means technically. You just have to be able to say “warm, late-afternoon light,” and let the app do the part I’d learned by trial and error.
The tiger and the mouse worked because I’d become fluent in a language I didn’t know I was learning. Mah-nah Mah-nah exists so nobody else has to spend their limited free queries learning it the same slow way.
Sandi is a Melbourne-based problem-solver, crisis-averter, and translator of the technical into the human. She spent decades being the person everyone called when something was broken, confusing, or just needed explaining properly — earning a reputation that preceded her wherever she went. Now she’s channelling that same instinct into AI: making it accessible, practical, and genuinely useful for people who think it isn’t for them.



