Models

Frontier models.
Production ready.

Access cutting-edge video, image, and text models - carefully selected,rigorously tested, and optimized for real production workflows at scale.

Carefully Tested

Hands-on evaluation before listing

Production Ready

Reliable quality and consistent outputs

Workflow Optimized

Works seamlessly with PomexAI workflows

We keep our catalog focused so you can find the right model faster.

PomexAI illustration

A focused catalog you can trust

Video
Seedance 2.5

Seedance 2.5

Text To VideoImage To VideoMulti-shot

Price

from $5.04/1M tokens

Max Duration

VideoMiniMax H3Media unavailable

MiniMax H3

MiniMax H3 is the MiniMax video route for text-to-video, image-to-video, and first/last-frame workflows, with audio generated natively on every clip. It supports 768P and 2K outputs at a fixed 24 fps, and is billed per second of generated video rather than per token.

Text To VideoImage To VideoMulti-shot

Max Duration

VideoMiniMax H3 MaxMedia unavailable

MiniMax H3 Max

Text To VideoImage To VideoMulti-shot

Max Duration

TextDeepSeek V4.1 Flash

DeepSeek V4.1 Flash

DeepSeek V4.1 Flash is a sparse mixture-of-experts model from DeepSeek, and the first built on the company's Causal Encoder-Decoder (CED) architecture. It activates 8B parameters on input and 16B on output from a 552B-parameter backbone, an asymmetric split that keeps per-token compute low relative to the model's total size. Image understanding is native to the architecture, with visual and text embeddings trained jointly from the start of pre-training rather than added afterward as in the earlier experimental V4 Flash Vision Exp(opens in new tab).

ChatReasoning

Price

$0.15/$0.6 per M

Context Window

VideoNEW
Seedance 2.0

Seedance 2.0

Currently ranked #1 in the Text-to-Video Arena; supports 12 multimodal reference inputs for high consistency.

Text To VideoImage To VideoMulti-shot

Price

from $2.4/1M tokens

Price detail

Max Duration

15s (High FPS)

Video
SkyReels V3

SkyReels V3

Add a multi-shot generation path for prompts, images, and reference-guided video that requires consistent motion and scene continuity.

Text To VideoImage To VideoMulti-shot

Price

$32 input / $32 output

Max Duration

18s

Our evaluation process

Four steps separate a research release from a production-ready model in our catalog.

1. Discover

Track releases from leading research labs, open-source communities, and commercial providers.

2. Benchmark

Test generation quality, inference speed, and API reliability against production requirements.

3. Stress Test

Build real workflows with the model to verify it performs under production constraints at scale.

4. Ship

Only models that consistently deliver production-quality results make it into the catalog.

Can't find the model you need?

Tell us what you're looking for. We're always evaluating new models.

Request a model