Sunday, October 11, 2026
Mobile Offer

🎁 You've Got 1 Reward Left

Check if your device is eligible for instant bonuses.

Unlock Now
Survey Cash

🧠 Discover the Simple Money Trick

This quick task could pay you today — no joke.

See It Now
Top Deals

📦 Top Freebies Available Near You

Get hot mobile rewards now. Limited time offers.

Get Started
Game Offer

🎮 Unlock Premium Game Packs

Boost your favorite game with hidden bonuses.

Claim Now
Money Offers

💸 Earn Instantly With This Task

No fees, no waiting — your earnings could be 1 click away.

Start Earning
Crypto Airdrop

🚀 Claim Free Crypto in Seconds

Register & grab real tokens now. Zero investment needed.

Get Tokens
Food Offers

🍔 Get Free Food Coupons

Claim your free fast food deals instantly.

Grab Coupons
VIP Offers

🎉 Join Our VIP Club

Access secret deals and daily giveaways.

Join Now
Mystery Offer

🎁 Mystery Gift Waiting for You

Click to reveal your surprise prize now!

Reveal Gift
App Bonus

📱 Download & Get Bonus

New apps giving out free rewards daily.

Download Now
Exclusive Deals

💎 Exclusive Offers Just for You

Unlock hidden discounts and perks.

Unlock Deals
Movie Offer

🎬 Watch Paid Movies Free

Stream your favorite flicks with no cost.

Watch Now
Prize Offer

🏆 Enter to Win Big Prizes

Join contests and win amazing rewards.

Enter Now
Life Hack

💡 Simple Life Hack to Save Cash

Try this now and watch your savings grow.

Learn More
Top Apps

📲 Top Apps Giving Gifts

Download & get rewards instantly.

Get Gifts
Summer Drinks

🍹 Summer Cocktails Recipes

Make refreshing drinks at home easily.

Get Recipes

Latest Posts

What is Decision 3.0? vLLM Semantic Router’s New Open Decision Models


The vLLM Semantic Router team has released Decision 3.0, a family of multimodal decision models. Decision 3.0 multimodal decision models read text, JSON and images, then answer typed questions about them. There are 5 sizes, from 0.8B to 27B, all under Apache-2.0. For developers, this is a fast way to classify, route and gate requests without parsing generated text.

TL;DR

  • Size: 5 models: d3-lite (0.85B), d3-nano (2.21B), d3-mini (4.54B), d3-flash (8.39B), d3 (26.09B). Context length: not disclosed.
  • Runs on: latency measured on 1 AMD Instinct MI325X GPU. BF16 weights; no official quantized variants listed.
  • Performance: the 27B d3 leads both the text and vision tables in its card, but on internal evaluation.
  • Best: 97.7 on KIE (CORD+FUNSD) document extraction (d3).
  • Worst: 34.1 on R-Bench-M (d3).
  • Bottom line:
    • Best: a 9B model now beats last generation’s 27B.
    • Worst: scores are self-reported against live board data.

What is Decision 3.0?

Decision 3.0 is a set of open decision models from vLLM Semantic Router that return probabilities instead of text. You pass a state (text or JSON), optional images, and named questions. Each question is a Choice, a Yes/No, or a Score on a scale. The model answers all questions in one call and returns a probability for every answer.

This is the ‘system one’ pattern for routers and guardrails. Instead of prompting an LLM and parsing its output, you get calibrated scores you can threshold.

How does d3 work?

Each model is a fine-tune of a Qwen base. The 27B d3 is built on Qwen3.8-27B. The smaller 4 build on Qwen3.5 checkpoints at 0.8B, 2B, 4B and 9B.

Every size carries a vision encoder: 0.10B in d3-lite, 0.33B in d3-nano and d3-mini, and 0.46B in d3-flash and d3. Requests accept several PNG, JPEG or WebP images, each read at up to 1.6 megapixels. Every question sees all images.

Each question gets its own forward pass over the input. Install targets transformers==5.17.0, and the optional flash-linear-attention package speeds up the linear-attention layers. Loading needs trust_remote_code=True.

How does Decision 3.0 perform on benchmarks?

Through the model card, the research team state the Jev Decision Index 0.3.1, a board that ranks open reproductions of TypeSafe’s Jev decision system. d3 numbers are the team’s internal evaluation.

On text, d3 scores 64.1, ahead of Perplexity Decider v1.1 (62.8) and Jev (60.1). One caveat: Torchcast Decision 27B posts a higher public-suite score (65.1 vs 64.9). On the vision board, d3 scores 71.6 against 70.6 for Perplexity Decider v1.1.

The size story is the strongest part. d3-flash (9B) scores 59.0, above Decision 2.0’s 27B at 55.9. On the public suite, gains over Decision 2.0 range from +8.0 (27B) to +15.1 (0.8B).

Class leadership is not uniform, though:

  • d3-mini (4B) and d3-flash (9B) lead both their text and vision tables.
  • d3-lite (0.8B) leads on text but ranks 3rd on vision (41.1, behind JPT-0.8B at 42.1).
  • d3-nano (2B) leads its vision table but trails LiquidAI d1-3B on text (35.8 vs 40.1).

On public vision tasks, d3 scores 97.3 on InfographicVQA and 89.9 on Mind2Web, both approximate rebuilds. It is weaker on MMMU-Pro vision (46.2) and Hateful Memes moderation (44.6).

How fast are the models?

Medians below are for 1 request at a time on 1 AMD Instinct MI325X, per each card:

  • d3-nano: 21.7 ms text, 92.0 ms with an image
  • d3-lite: 23.5 ms text, 60.7 ms with an image
  • d3-mini: 28.5 ms text, 115.0 ms with an image
  • d3-flash: 29.7 ms text, 171.8 ms with an image
  • d3: 84 ms text, 342 ms with an image

All models were trained on AMD Instinct MI325X GPUs.

How does d3 compare with other decision models?

Feature d3 (vLLM SR) Perplexity Decider v1.1 Decision 2.0 Vega-27B LiquidAI d1-3B
Parameters 26.09B 26B (HF listing) 29.37B 3.12B
Base model Qwen3.8-27B Qwen3.8-27B Qwen3.8-27B LFM2.5-VL-3B
Inputs Text, JSON, images Text, images Text, JSON Text, JSON, images
Context Not disclosed 8,192 tokens per decision 32,768 tokens 32,768 tokens
License Apache-2.0 Apache-2.0 Apache-2.0 LFM 1.0
Decision Index (own card) 64.1 (v0.3.1, internal) 61.56 (version not stated) 56.5 (v0.2.1) 48.57 (v0.2.1)
Index 0.3.1 board (per d3 cards) 64.1 62.8 55.9 40.1
Vision board 0.3.1 71.6 70.6 Not disclosed (text only) Not disclosed
Median latency (own card) 84 ms text, 1 MI325X Not disclosed 71.4 ms, 1 GPU 9 ms, 1 MI325X

Key Takeaways

  • 5 Apache-2.0 decision models, 0.8B to 27B, all with image input.
  • d3 (27B) tops its text (64.1) and vision (71.6) tables, on internal evals.
  • The 9B d3-flash (59.0) beats Decision 2.0’s 27B (55.9).
  • The 0.8B and 2B models do not lead every board.
  • Median text latency runs 21.7 to 84 ms on 1 MI325X.

Check out the Hugging Face collection, the d3 model card and the vLLM Semantic Router GitHub repo. All credit goes to the researcher of this project. Also, feel free to follow us on Twitter and don’t forget to join our 150k+ML SubReddit and Subscribe to our Newsletter. Wait! are you on telegram? now you can join us on telegram as well.


Michal Sutter is a data science professional with a Master of Science in Data Science from the University of Padova. With a solid foundation in statistical analysis, machine learning, and data engineering, Michal excels at transforming complex datasets into actionable insights.



Source link

Latest Posts

Don't Miss

Stay in touch

To be updated with all the latest news, offers and special announcements.