
Requires IMA API Key, one-stop creation for images, videos, music, and speech
IMA-all-ai is a comprehensive, multi-modal AI generation and media orchestration skill on EasyClaw. It integrates directly with Tencent's IMA open API endpoints to automate the creation of high-quality digital assets across four core media formats — including text-to-image graphic design, image-to-video promotional rendering, text-to-music ambient background generation, and natural text-to-speech audio conversions — through simple, conversational commands, requiring only a single authorized API key.
The skill is built for e-commerce marketers creating multi-media ad campaigns, content creators preparing video soundtracks, and developers building interactive, media-rich web applications.
The expected outcome is a compiled, production-ready media asset link (image, video, audio) returned directly in your conversation, ready to download and deploy on your marketing channels.
1. API configuration and verification. The skill establishes a secure HTTPS connection with Tencent's IMA servers using your private developer Client ID and API Key, validating active credit balances before execution.
2. Text-to-Image rendering. Provide your descriptive prompt (e.g., "draw a cyberpunk cat with neon lights") and specify the aspect ratio (square, portrait, or landscape). The skill submits the job and retrieves the generated image link.
3. Image-to-Video compilation. For video requests, provide a product photo or reference image and describe the motion (e.g., "turn this photo into a 10-second promotional video"). The skill runs the video compilation pipeline headlessly.
4. Text-to-Music generation. To generate background tracks, describe the genre and mood (e.g., "create a lo-fi background track for my vlog"). The skill's music generator synthesizes the audio, returning a downloadable MP3 link.
5. Vocal speech synthesis. For text-to-speech tasks, the skill converts your written script into natural-sounding voiceover audio, supporting diverse character profiles and accent tunings.
- Multi-modal generation cockpit: Generates images, videos, music, and speech from a single interface.
- High-fidelity text-to-image: Renders professional graphics and illustrations from raw prompts.
- Image-to-Video rendering: Compiles 10-second promotional videos from static product photos.
- Text-to-Music synthesizer: Generates high-quality ambient and background MP3 music tracks.
- Natural speech converter: Converts text scripts to natural-sounding voiceover files headlessly.
- Secure token management: Handles private credentials and IMA keys securely in the background.
1. Creating a cohesive multi-media ad campaign
A growth marketer wants to launch a campaign for a new coffee brand. Instead of hiring separate graphic designers, video editors, and audio engineers, they use the skill: first generating a stylized product image ("draw a warm, steaming coffee cup on a rustic wood table"), then compiling a 10-second promotional video from the image, and finally generating a relaxed lo-fi background music track, creating a complete ad asset package in minutes.
2. Animating a static product photo for social ads
An e-commerce seller has a high-quality product photo of an insulated water bottle. To stand out on social feeds, they require video. They upload the photo to the workspace. The skill processes the asset: animating the background, adding realistic water splash transitions, and outputting a professional, 10-second MP4 promo video.
3. Generating a custom background track for a video blog
A vlog creator needs a unique, copyright-free background track for their daily tech review. They describe the vibe: "create a mellow, ambient electronic track with soft synth pads." The skill connects to IMA's music engine, synthesizes the track, and delivers a downloadable, high-fidelity MP3 link, bypassing copyright strikes.
4. Converting scripts to voiceovers for tutorials
A technical writer wants to add voice narration to a software tutorial video. They provide the written script. The skill converts the text into a natural-sounding, clear male English voiceover file, ready to sync with their video timeline.
5. Generating custom illustrations for presentations
An analyst wants to add a specific visual metaphor to their presentation slides (e.g., "a rocket blasting through data clouds"). The skill renders a high-quality, digital-vector style illustration from the prompt, saving hours of searching through generic stock image sites.
A marketer wants to generate a high-quality image and matching background track for a product promo.
1. They open EasyClaw and activate IMA-all-ai.
2. They run: *"Draw a cyberpunk cat with neon lights, 16:9, and generate a 10-second matching ambient track."*
3. The skill connects to the IMA API, submits the text-to-image job, and fetches the rendered PNG.
4. It calls the music engine to synthesize the ambient audio, returning the MP3 file.
5. It outputs both the image download link and the audio file link in the conversation.
Cohesive multi-media assets generated in under 90 seconds.
Add this skill to your EasyClaw workspace
Describe your task in a chat message
Review the output and iterate if needed
Export or share the results directly from EasyClaw
Combine with other skills to build automated workflows
Yes. To generate images, videos, music, or speech, you must configure a valid IMA Client ID and Secret key in your EasyClaw environment variables. The skill will guide you through the setup.
The image-to-video compilation engine is highly optimized to generate high-impact, 10-second promotional and social media ad video clips.
The music generator synthesizes high-fidelity audio, delivering completed tracks as standard, highly compatible MP3 download links ready for any media editor.
The generated images, videos, and music are completely original, synthesised from your prompts, and are free of copyright restrictions, allowing you to deploy them safely in commercial campaigns.
The image generator natively supports: Square (1:1), Landscape (16:9 for YouTube), Portrait (9:16 for TikTok/Reels), and standard blog ratios (4:3).
Video rendering is a heavy, compute-intensive task. It typically takes 30 seconds to 2 minutes depending on complexity. The skill polls the server status asynchronously and alerts you once the final download link is ready.
Yes. The speech converter supports multiple voice profiles, including male, female, professional narrators, and child characters, across English and Chinese dialects.
Yes. All data processing, prompt parsing, and asset deliveries are executed directly between your local machine and Tencent's secure servers over HTTPS, keeping your intellectual property private.
Yes. You can provide a reference image URL and describe your changes (e.g., "change the neon lights from blue to red") and the image engine will modify the asset accordingly.
All generated images, videos, and audio files are logged locally under `public/data/media_exports/` in your workspace exports directory, paired with a direct local link in your chat.
Browse more in General Tools or all skills.
Get EasyClaw, add this skill, and start building AI agent workflows in minutes.
Get EasyClaw Free →