Google Veo 3 & Gemini Image MCP Server for Claude
We’ve released an open-source MCP server that brings Google’s latest visual AI models directly into Claude or any MCP-compatible AI agent. With Gemini 3 Pro Image (Nano Banana Pro), Imagen 4, and VEO 3, AI agents can now generate studio-quality images and cinematic videos with synchronized audio through natural conversation.
Available Now on Adomo.ai
Try it live on Adomo.ai, our enterprise AI platform, where Gemini Media MCP runs with no setup.
For developers and teams who want to run it themselves, we’ve made it available across multiple platforms:
- Docker Hub, pull and run in seconds
- GitHub, full source code and documentation
- PyPI, install via pip
Real Business Impact
It’s built for production work:
E-commerce
Generate product shots and promotional videos at scale while maintaining brand consistency. No more waiting on design teams for every product variation.
Marketing
Create complete campaign assets, from static images to video ads, without leaving your AI workflow. Iterate faster and launch campaigns sooner.
Enterprise
Produce training materials, technical diagrams, and localized content across markets instantly. Scale your content operations without scaling your team.
Why We Open-Sourced This
We originally built Gemini Media MCP for Adomo AI’s enterprise platform. Pairing Gemini’s image understanding with VEO’s video generation worked well enough that keeping it internal made little sense.
Every AI agent should be able to create professional visual content. We’re committed to continuing to release and support open-source AI tools for the community.
Created with Claude and Gemini Media MCP
The image and video below were generated using Claude with the Gemini Media MCP server:

Get Started
The server works with:
- Claude Desktop
- Claude Code
- Any MCP-compatible platform
Head over to GitHub to get started, or experience it directly on Adomo.ai.
We’d like to hear your feedback. Open an issue on GitHub or reach out to us directly.
Thanks to the teams at Google and DeepMind for building these models.