How the model works

Draft Transformer · Dota 2 neural network

How the modelevaluates your draft.

Draft Transformer estimates win probability from both lineups and match settings. It learns from played matches, and the drafter compares its predictions to help you choose your next hero.

01 / Hero map

How the modelrepresents heroes.

The model describes each hero with a set of numbers called an embedding. It learns these descriptions from matches. Select a hero on the map to explore other heroes with similar representations.

Map preview

Try searching and selecting heroes. Positions and links are examples.

1.0×
Hero map127 heroes
Anti-MageAxeBaneBloodseekerCrystal MaidenDrow RangerEarthshakerJuggernautMiranaMorphlingShadow FiendPhantom LancerPuckPudgeRazorSand KingStorm SpiritSvenTinyVengeful SpiritWindrangerZeusKunkkaLinaLionShadow ShamanSlardarTidehunterWitch DoctorLichRikiEnigmaTinkerSniperNecrophosWarlockBeastmasterQueen of PainVenomancerFaceless VoidWraith KingDeath ProphetPugnaTemplar AssassinViperLunaDragon KnightDazzleClockwerkLeshracNature's ProphetLifestealerDark SeerClinkzOmniknightEnchantressHuskarNight StalkerBroodmotherBounty HunterWeaverJakiroBatriderChenSpectreAncient ApparitionDoomUrsaSpirit BreakerGyrocopterAlchemistInvokerSilencerOutworld DestroyerLycanBrewmasterShadow DemonLone DruidChaos KnightMeepoTreant ProtectorOgre MagiUndyingRubickDisruptorNyx AssassinNaga SirenKeeper of the LightIoVisageSlarkMedusaTroll WarlordCentaur WarrunnerMagnusTimbersawBristlebackTuskSkywrath MageAbaddonElder TitanLegion CommanderTechiesEmber SpiritEarth SpiritUnderlordTerrorbladePhoenixOracleWinter WyvernArc WardenMonkey KingDark WillowPangolierGrimstrokeHoodwinkVoid SpiritSnapfireMarsRingmasterDawnbreakerMarciPrimal BeastMuertaKezLargoPhantom Assassin
Selected hero

Phantom Assassin

Heroes in this example 5

How to read this map

Select a point, search for a hero or choose a name in the list to explore its neighborhood. Preview positions and links let you try the interface; they do not represent model results.

Use + / − to zoom. The wheel scrolls the page.

02 / From heroes to combinations

The same hero.A different valuein each lineup.

The embedding is a starting point. The model then considers allies, opponents, roles, lanes and match conditions. A hero’s representation is updated with this surrounding information.

Why a Transformer

Attention layers pass information between picks. Later layers work with representations that already contain context. Together with nonlinear transformations, this allows the model to learn effects where one hero’s contribution depends on several others at once.

Same pair. Different situation.

Add heroes and see what changes

Looking at just two heroes for now

Sniper is within reach

Phantom Strike lets PA close the distance to Sniper. Looking at this pair, the plan is clear: find the target and get into melee range.

A plan to reach the target
+ Abaddon and Axe on Sniper’s side

The jump becomes a risk

Add Abaddon and Axe. A shield helps Sniper survive damage, while Axe can catch PA after her jump. The pair is unchanged, but securing the kill is a different problem.

Jumping into saves and control
+ Tidehunter on PA’s side

Initiation changes the opening

Add Tidehunter alongside PA. If Ravage catches Sniper’s protection, PA can follow up. Part of Tide’s value here comes from helping another hero execute their plan.

A team plan becomes possible
Illustrative gameplay scenario. Lines show hero interactions, not attention weights or model results.

The model receives match outcomes, not predefined rules about saves or counter picks. The gameplay example illustrates context; it does not explain a particular prediction.

03 / Training

Learning fromplayed matches.

The model compares its prediction with the match result and adjusts its internal parameters, or weights. Some heroes, roles and lanes are hidden during training so it also learns to work with incomplete drafts.

One outcome. Different amounts of information.

Explore the inputs the model receives during training

Model input10 / 10 heroes visible
Radiant
  1. Phantom AssassinRole: CoreLane: Safe
  2. InvokerRole: CoreLane: Mid
  3. TidehunterRole: CoreLane: Off
  4. MiranaRole: SupportLane: Off
  5. Crystal MaidenRole: SupportLane: Safe
Dire
  1. JuggernautRole: CoreLane: Safe
  2. SniperRole: CoreLane: Mid
  3. AxeRole: CoreLane: Off
  4. LionRole: SupportLane: Off
  5. AbaddonRole: SupportLane: Safe
Every hero is visible

The inputs include heroes, roles and lanes. Training connects this information with the outcome of the match.

Model input5 / 10 heroes visible
Radiant
  1. Phantom AssassinRole: CoreLane: Safe
  2. InvokerRole: CoreLane: Mid
  3. TidehunterRole: CoreLane: Off
  4. Hidden heroRole: SupportLane: Off
  5. Hidden heroRole: SupportLane: Safe
Dire
  1. JuggernautRole: CoreLane: Safe
  2. SniperRole: CoreLane: Mid
  3. Hidden heroRole: CoreLane: Off
  4. Hidden heroRole: SupportLane: Off
  5. Hidden heroRole: SupportLane: Safe
Part of the lineup is hidden

This illustration keeps 3 Radiant heroes and 2 Dire heroes visible. A slot can retain its role and lane even when its hero is hidden.

Model input10 / 10 heroes visible
Radiant
  1. Phantom AssassinRole: CoreLane: Safe
  2. InvokerRole: CoreLane: hidden
  3. TidehunterRole: CoreLane: Off
  4. MiranaRole: hiddenLane: Off
  5. Crystal MaidenRole: SupportLane: Safe
Dire
  1. JuggernautRole: CoreLane: Safe
  2. SniperRole: hiddenLane: Mid
  3. AxeRole: CoreLane: Off
  4. LionRole: SupportLane: Off
  5. AbaddonRole: SupportLane: hidden
Some details are missing

All heroes remain visible, but some slots lose a role or lane. These features are hidden independently of each other.

Training labelRadiant win

Stays the same as the inputs change

An illustration of preparing an example. This is not a real match or a model prediction.A masking phase is sampled for 40% of training examples; the possible phases include a full draft. Roles and lanes are hidden independently, with a 10% probability per feature. Match duration is always hidden during validation.
  1. 01

    Receive lineups and conditions

    Heroes, sides, roles and lanes. Patch, mode, lobby type and average rank provide context.

  2. 02

    Work with hidden inputs

    Some heroes may be unknown. Roles and lanes are hidden independently; the match outcome stays the same.

  3. 03

    Adjust the prediction

    The loss function compares the estimate with the result. Training gradually updates representations and connections.

04 / Measuring quality

How we checkprediction quality.

The model learns from 95% of matches. On the remaining 5%, we compare its predictions with actual outcomes. Errors on these validation matches do not update model parameters; the results help us select the best saved version.

95%Training matches
5%Validation matches

Matches are split at random while keeping the same patch proportions in both groups. This checks performance on represented patches. A new patch needs a separate evaluation.

Duration is hidden

Validation does not reveal when the match ended. The drafter’s main prediction also uses no specified duration.

Comparing model versions

We select the version with the best validation ROC AUC. Since these matches help select the version, an independent final test needs another dataset. Separate results for each draft stage are not available yet.

Validation results

Metrics have not been published yet

ROC AUC—
How often winning lineups receive a higher estimate than losing ones. Closer to 1 is better; 0.5 corresponds to random ranking.
Accuracy—
The share of matches won by the team given more than a 50% chance. This measures predicted winners, not your win rate.
LogLoss—
Accounts for confidence: a wrong prediction at 90% is penalized more than one at 55%. Lower is better.

Measured results will appear here when published. Until then, the architecture and hero map alone cannot tell you how accurate the model is.

Architecture and training parameters

Self-attention and nonlinear SwiGLU blocks form the core. Layer count does not specify a fixed order of gameplay interactions: dependencies develop through training.

Architecture
4 Transformer blocks · 10 attention heads
Representations
256 hero features + role and lane
Attention
QK-Norm · Scaled Dot-Product Attention
Inside each block
RMSNorm · SwiGLU · residual connections
Regularization
Dropout · label smoothing · input masking
Optimization
PyTorch · Lightning · AdamW

Attention Is All You Need

05 / How the drafter works

What changes if youpick another hero?

The drafter inserts available heroes into a selected slot and compares predictions while keeping the other picks fixed. This lets you assess options from your own pool in a specific lineup.

  1. 01

    Set up the draft

    Add known picks and choose the match conditions.

  2. 02

    The model compares replacements

    Every candidate is evaluated in the same surrounding lineup.

  3. 03

    Choose from your pool

    Consider the recommendations alongside your experience and plan.

The model evaluates the lineup. Your hero experience, team coordination, items and in-game decisions remain outside the prediction.

Understanding predictions

Before youtrust the estimate

What does the hero embedding atlas show?

It projects base hero representations before processing a particular draft. Similar vectors do not necessarily imply synergy, a counter pick or identical play styles. t-SNE simplifies high-dimensional data; neighbors are calculated separately using original vectors. In preview mode, positions and links are illustrative.

How is the neural network different from a Dota 2 counter-pick table?

A pairwise table describes two heroes. Draft Transformer receives entire lineups and their context. The architecture can account for combinations where one pick’s effect depends on several others. How successfully it learns these dependencies needs to be evaluated on data.

Can I compare heroes before a draft is complete?

Yes. Training includes examples with hidden heroes. The drafter can compare options while slots remain empty. Every new pick adds context and changes the calculation. Separate quality metrics for each draft stage have not been published yet.

Is win probability the model’s accuracy?

No. A probability refers to one lineup and its conditions. Accuracy is the share of correctly predicted outcomes across validation matches. Even a high estimate for a draft does not guarantee a win.

Does the time chart predict how a match unfolds?

No. It compares estimates conditioned on specified match durations. It is neither a minute-by-minute simulation nor a prediction of when your game will end. The main prediction uses unknown duration.

How can I try the model on my match?

Open the online drafter, add heroes manually or import a public match by ID. Compare next-pick options and lineup estimates. The drafter is free and requires no account.

Try it with your lineup

Find a herofor your next match.

Build a draft and compare a few options. See how win probability changes with different allies and opponents.