Making images that say what you mean | Master AI Automation in 4 hours Master AI Automation in 4 hours Course About Ayush Modules Sample chapter Toolbox The Microcap Minute Classroom / Module 09: Images, Voice & Video / Chapter 1 Making images that say what you mean Watch first, then read. Same lesson, your pace. What you will learn – The four-part image prompt – Where to generate for free – Text-in-images and the fixes The four parts Image models predict pictures from descriptions, so description quality rules results. Structure prompts as: Subject , what , specifically: “a Bengal cat wearing sunglasses” not “a cool cat” Style , how rendered : “watercolour illustration”, “3D render”, “1990s textbook diagram”, “studio photograph” Composition , framing : “close-up”, “wide shot from above”, “centred, plain background” Constraints , “no text”, “muted colours”, “white background” “A friendly cartoon tiger explaining fractions at a blackboard, flat vector illustration, centred, white background, no text.” Iterate like Module 2 taught, one change per generation. And keep a prompts/images.md of winners; image prompts are even more re-usable than text ones. Free-first options ChatGPT’s image generation (free tier includes limited generations), strong all-rounder, good with instructions Gemini’s image generation (free tier), fast iteration, decent quality Bing/Microsoft Image Creator , free DALL·E-based generations with daily boosts Ideogram / similar freemium tools , notably better at rendering readable text inside images Paid leaders (Midjourney etc.) earn their price in aesthetics and control, after you’ve exhausted free learning capacity, not before. Text in images: the honest warning Models historically mangle words inside images (“STUDY SESION”). Fixes: put text in afterward with any editor, use text-specialised tools (Ideogram-class), or accept stylised gibberish where it doesn’t matter. Check spelling on every generated poster before showing anyone. One more habit from Module 11 (preview): generated photorealistic people and events carry ethical weight. Cartoon diagrams of tigers don’t. Keep the line visible. Try it yourself Generate the tiger-fractions prompt above in two different free tools. Then run a controlled experiment: same subject + style, three different compositions; note how framing changes usefulness. Finally create one real asset you need, a diagram for notes, an event poster, a thumbnail, iterating until it says what you mean . Save prompt + final into learn/image-lab.md . Key takeaways – Image prompts = subject + style + composition + constraints; iterate one variable at a time. – Free tiers cover learning completely; paid buys polish later. – In-image text is unreliable, place text afterwards or use specialised tools. – Photorealistic people/events need Module 11’s honesty rules. Download the exercise sheet (PDF) Module workbook (PDF) ← Module 09 index Next: Vision: making AI see → Classroom / Module 09: Images, Voice & Video / Chapter 1 Making images that say what you mean What you will learn – The four-part image prompt – Where to generate for free – Text-in-images and the fixes The four parts Image models predict pictures from descriptions, so description quality rules results. Structure prompts as: Subject , what , specifically: “a Bengal cat wearing sunglasses” not “a cool cat” Style , how rendered : “watercolour illustration”, “3D render”, “1990s textbook diagram”, “studio photograph” Composition , framing : “close-up”, “wide shot from above”, “centred, plain background” Constraints , “no text”, “muted colours”, “white background” “A friendly cartoon tiger explaining fractions at a blackboard, flat vector illustration, centred, white background, no text.” Iterate like Module 2 taught, one change per generation. And keep a prompts/images.md of winners; image prompts are even more re-usable than text ones. Free-first options ChatGPT’s image generation (free tier includes limited generations), strong all-ro
Making images that say what you mean
Written by
in