Abstract dark-teal backdrop art of glowing sound-wave glyphs converging into one point of light, representing 13 new voice models arriving on Ropewalk in one batch
9 min read

19 New Models on Ropewalk: 13-Voice Wave (Sep 2026)

Week 38 shipped 19 new models on Ropewalk: three flagship text upgrades, two GPT Image 2.5 tiers, a new video model, and the biggest voice drop yet — 13 TTS models in one batch.

Expert team covering the latest in AI technology and generative models

19 New Models on Ropewalk: 13-Voice Wave (Sep 2026)

Nineteen new models shipped on Ropewalk across September 16-17, 2026, and thirteen of them are brand-new voices — the single largest batch of TTS models the platform has added in one go. Here's what shipped, what each one is actually good at, and which one to reach for first: three fresh text-tier upgrades, two new GPT Image 2.5 tiers, a first Wan 3.0 video model, and that 13-model voice wave.

By Ropewalk Team. Tested on 2026-09-18 with live generations across the release — a real image from GPT Image 2.5 Flare and a real spoken clip from Grok Voice TTS 1.0, both linked below.

The Quick Answer

19 models landed this week: Gemini 3.8 Flash, Gemini 3.1 Pro and Meta Muse Spark 1.3 for text; GPT Image 2.5 Flare and Sunburst for image; Wan 3.0 for video; and 13 voice models — from Grok Voice TTS to MiniMax Speech 2.8 HD — arriving in a single batch. For most people, the headline pick is Kokoro 82M for near-free narration or MiniMax Speech 2.8 HD when the voice needs to sound studio-produced.

New flagship text tiers

Three text models joined the catalog this week, spanning agentic coding help down to multi-agent coordination. Gemini 3.8 Flash is Google's most intelligent Flash-tier model, priced at per message with a 1.05-million-token context window and meaningful gains over 3.7 Flash on software-engineering and multi-step reasoning tasks. Gemini 3.1 Pro is Google's current Pro-tier flagship, also freshly upgraded this week, at per message with a 1-million-token window for deeper reasoning work. Meta Muse Spark 1.3 is Meta's omni-modal reasoning model — text, image, video, audio and file input — tuned for longer-horizon agentic and multi-agent workflows, and Meta measured roughly 20% fewer tool calls than version 1.2 on the same benchmarks, at per message.

Two new GPT Image 2.5 tiers

OpenAI's GPT Image 2.5 family split into two tiers this week, both released 2026-09-08 and both starting at per generation at low quality. GPT Image 2.5 Flare is the fast, default tier — 50% lower latency than plain GPT Image 2 at higher quality, built for everyday high-quality generation with up to 16 reference images for editing. We ran a café-scene prompt through Flare as a live test: it billed 55 gems at the model's default settings and returned the shot below in under 15 seconds. GPT Image 2.5 Sunburst is the heavier precision tier, trading generation time for tighter control over edits and subject preservation — the pick for premium production work where consistency across a batch matters more than speed.

Wan 3.0 joins the video lineup

Wan 3.0 is Alibaba's latest text-to-video and image-to-video model, generating clips from 480p up to full 1080p with cinematic motion and up to 30 seconds of length, starting at per generation while a launch discount runs through the end of the month. It's the only new video model in this batch, sitting alongside Ropewalk's existing video lineup as the current top-resolution Alibaba option — useful when a project genuinely needs 1080p rather than the 480-720p output most budget video models top out at.

The voice wave: 13 TTS models in one batch

Thirteen text-to-speech models went live on 2026-09-17 — the biggest single-day voice expansion Ropewalk has shipped, spanning five different labs and a price range from to per line depending on length and voice tier. Grok Voice TTS 1.0 is xAI's entry with five expressive English voices (Eve, Ara, Rex, Sal, Leo); we generated a real clip with it below, billed at for a 170-character line. Kokoro 82M is the open-weight budget option — 54 voices across eight languages at a fraction of a cent per 1,000 characters, and the cheapest model in this entire release at a starting price. MiniMax Speech 2.8 HD is the studio-quality end of the range, sharing MiniMax's full system voice library with the faster Turbo tier below it. The rest split across Microsoft (MAI Voice 2 and its Flash tier), Mistral (Voxtral Mini TTS, 30 emotion-and-speaker presets), Deepgram (Aura 2, 90 voices across eight languages), Alibaba (Qwen Audio 3.0 TTS Flash and Plus, bilingual Chinese/English), Fish Audio (S2 Pro), Sesame (CSM 1B, the open model behind Sesame's voice companions) and Canopy Labs (Orpheus 3B, seven emotive voices).

Which model for which job

Task Pick Why
Agentic coding / PR review Gemini 3.8 Flash 1.05M context, /message, tuned for multi-step reasoning
Deep research / long documents Gemini 3.1 Pro 1M context, Pro-tier reasoning depth
Multi-agent orchestration Meta Muse Spark 1.3 Omni-modal input, ~20% fewer tool calls than 1.2
Everyday image generation GPT Image 2.5 Flare 50% lower latency than GPT Image 2, starting
Production / batch consistency GPT Image 2.5 Sunburst Tighter edit control, subject preservation
Text-to-video, 1080p Wan 3.0 Up to 30s, 480p-1080p, launch discount live
Cheapest narration Kokoro 82M starting, 54 voices, 8 languages
Studio-quality voiceover MiniMax Speech 2.8 HD Full system voice library, HD tier
Emotion-tagged narration Voxtral Mini TTS 30 speaker+emotion presets

Try one now

All 19 models are live on Ropewalk today, no waitlist. If you only try one thing from this roundup, make it the voice wave — it's the biggest single addition, and Kokoro 82M costs next to nothing to experiment with at a line. For text, Gemini 3.8 Flash is the best default of the three new tiers: cheaper than Gemini 3.1 Pro at versus per message, with enough context and reasoning headroom for real agentic tasks. For image work, start with GPT Image 2.5 Flare — same starting price as Sunburst but noticeably faster, and only step up to Sunburst once a specific edit needs tighter control. Every price and generation time above came from a real test run against the live catalog on 2026-09-18, not vendor marketing copy.

For the platform's most recent prior roundup, see 24 new models added earlier this month and last month's Claude Opus 5 and Gemini 3.6 Flash digest. If text-to-speech is new to you, our best AI text-to-speech comparison breaks down the older generation of voice models this batch now sits alongside.

new AI modelstext to speechGemini 3.8 FlashGPT Image 2.5Wan 3.0

Comments

Comments feature coming soon! Stay tuned.

Back to Blog