Source Models
The family of specialised models that flag interesting players — Breakout, Hidden Value, Undervalued, Peak Value, Sell-Fee and Risers — each tuned to one recruitment question and validated on its own outcome, unified behind a single flag contract the product reads.
A Source Model is a self-contained model that answers one scouting question — who is about to break out? who is hidden in a lower league? who is mispriced? — and writes its answer as a flag on a player. Several very different models all speak the same contract, so they surface side by side as the filterable chips and badges you see on the discovery screens, each one telling you why a player is worth a second look.
In one line: there is no single "good player" detector. Recruitment is many questions that genuinely conflict, so we run many specialised models — each tuned to one question and validated on its own outcome — and unify their outputs behind one contract the product reads without caring how any individual model works.
Why a family, not one model
A single score can only express one notion of "interesting", and recruitment has several that pull in different directions:
- Will this teenager reach an elite league? is a five-year forecast.
- Is this player about to jump in value? is a near-term breakout signal.
- Do his underlying numbers outrank his reputation? is a point-in-time mispricing.
- How high could his value go, and what would he sell for? are valuation questions.
Each wants a different model, a different group of comparable players, a different label, and a different time horizon. Forcing them into one number would average away exactly the disagreements that make a player worth a closer look. So each question gets its own model, and the product standardises on the contract — a confidence score, a trajectory shape, data-driven comparables — not on the model. A new model can be added without touching the screens: emit a flag in the agreed shape and it appears.
The models
Breakout — whose recent form looks like a player's right before a big jump in value? A gradient-boosted classifier scores each player's recent match-by-match output, then that raw score is empirically calibrated against history: we look up how often players in the same league, age and score band actually went on to break out, so the headline number reads as a real rate rather than an uncalibrated score.
Hidden Value — which lower-league youngster will reach an elite league? Built on the career-reading forecaster, run as an ensemble, it reports the probability of reaching a top-two European league within five years — and shows the ensemble's internal disagreement as an honest measure of how sure it is.
Undervalued — whose underlying numbers outrank his headline rating? The simplest model, pointedly so: the gap is the signal. Within a player's position-and-league group we rank his underlying performance and his headline match rating; a player who ranks far higher on substance than on reputation is flagged. Because there's no forward label, this one is descriptive — it states a present mispricing, it doesn't forecast a move.
Peak Value — how high could his value go, relative to today's price? A gradient-boosted model predicts a player's career-peak value, expressed as a within-cohort rank so it's comparable across ages and positions. An upside filter drops players already at their peak, so it surfaces genuine room to grow rather than just re-listing the names everyone knows.
Sell-Fee — who would command a strong sale fee? — for planning outgoings and renewals. A gradient-boosted model ranks players within their position-and-age cohort. We deliberately withhold the euro figure here: the model ranks players reliably but isn't yet calibrated on the absolute number, so only the trustworthy rank is shown.
Risers — which under-26 player is better than his club context implies? — the product surface of the Player Rating system: the players whose ability runs furthest ahead of their club, age and division.
How we keep it honest
Each model is validated on its own outcome, out-of-sample wherever a forward label exists:
- Breakout — on a rolling monthly hold-out, it identifies imminent breakouts with strong precision at the top of its list, well above the base rate.
- Peak Value and Sell-Fee — rank players in close agreement with what actually happened, beating the naive "today's value" baseline; Peak Value is strongest in the mid-20s and weakest for the very youngest, which is why it focuses on under-28s.
- Hidden Value — validated against players' realised five-year outcomes, with ensemble disagreement surfaced rather than hidden.
- Undervalued — has no forward label by design; it is a descriptive mispricing measure, not a prediction.
A confidence score is only comparable within a model, never across them — a Breakout 80 and a Sell-Fee 80 mean different things, so the product ranks within each model and we say so.
What it can't do
- Scores aren't cross-model comparable. Each model derives its confidence differently; rank within a model, don't compare across them.
- Some euro figures are withheld until they're calibrated well enough to trust.
- Undervalued can lag. With no forward label, it can flag a player whose rating is merely catching up to noisy underlying numbers.
- Comparables need history. "Resembles this player at the same age" needs an encodable career; a flag without comparables is a coverage gap, not a model failure.
- Coverage, not model. Flags only exist for the leagues we cover; and each flag is a point-in-time statement refreshed daily, not a continuously updated value.
The research behind it
The models draw on gradient boosting and probability calibration, the football market-value literature, the sports-economics framing of mispricing, and multi-task forecasting.
- Friedman, J. H. (2001). Greedy function approximation: a gradient boosting machine. Annals of Statistics. — The gradient-boosting framework behind the classifiers and valuation models.
- Ke, G. et al. (2017). LightGBM: A Highly Efficient Gradient Boosting Decision Tree. NeurIPS. — The efficient learner used for the boosted models.
- Zadrozny, B. & Elkan, C. (2001). Obtaining calibrated probability estimates… ICML. — The empirical calibration that turns a raw breakout score into a real rate.
- Müller, O., Simons, A. & Weinmann, M. (2017). Beyond crowd judgments: Data-driven estimation of market value in association football. European Journal of Operational Research. — The reference for the valuation models.
- Hakes, J. K. & Sauer, R. D. (2006). An economic evaluation of the Moneyball hypothesis. Journal of Economic Perspectives. — The market misprices specific skills; the framing behind Undervalued and Risers.
- Kendall, A., Gal, Y. & Cipolla, R. (2018). Multi-task learning using uncertainty to weigh losses. CVPR. — The multi-task training behind the shared forecaster.
Keep reading
- Player Rating
- Club & League Strength
- Market Values
- Cross-League Projection
- Trajectory & Ceiling Forecasting
- Career Outlook
- Recommendations
- Tactical & Realistic Fit
- Connection Degree
- Work-Permit Eligibility
- Identity Resolution
- How the Data Stays Correct
Oraca is in private beta with a small number of clubs.
Request early access