Ken Barry · Cork · sports trading systems since 2012 · this repository since April 2016

Models that beat the bookmakers.

Lure is the e-sports prediction platform I have built and run alone since 2016. Its heart is machine learning: Counter-Strike models, in use since 2017, that beat the bookmakers for years until around twenty of them had banned my accounts, and now models for Dota 2, League of Legends, Valorant, StarCraft 2, Mobile Legends and tennis. It pulls match data and prices from forty-five sources, reconciles them into one event graph, and retrains and redeploys every model every night without a human. lure.ie is not the product and does no trading: it is the read-only window I built so I could check on the engine while travelling, and it is switched on so you can look in.

5,718commits since April 2016
222,000lines of Java, written alone
45price and data feeds
1,845engineered features, CS map model

What it is

A model is only as good as its idea of the event.

Forty-five venues and data providers each have their own idea of what a match is called, when it starts, which market is which and which side is which. Before any model can be measured against the market, Lure has to decide that two rows from two different books are the same runner in the same market of the same event, and be right every time.

Everything downstream depends on that. A model is only being measured against the market if the runner it priced is the runner the bookmaker is quoting, and a bet is only the bet the model meant if the mapping is correct. Most of the code in this repository exists to earn that identity and then keep it consistent while forty-five sources disagree in real time.

It runs continuously on one Windows machine in Cork, backed by MySQL, with a read-only monitoring window relayed out to lure.ie. I designed, wrote, deployed and operate all of it. Across 6,770 commits on the two repositories, thirteen carry anyone's name but mine — recent ones stamped by a coding assistant working on my machine.

lure.ie with one Counter-Strike event opened, baks versus spirit academy at 15:00: the Match Winner market header reads 101.00 percent with the model's 102.22 percent beside it, followed by collapsed groups for Match Handicaps, Map 1 to Map 3 markets, Map Selections and Specials, and the rest of the day's fixtures below, each row coloured by the model's read and carrying its source-coverage icons
One event opened on lure.ie this afternoon. Every row is a canonical event reconciled from whichever of the forty-five sources carry it, and the icons on the row say which ones do. Inside the event, each market group is the same reconciliation one level down: the Match Winner header shows the book at 101.00% next to the model's own 102.22%, and the row colour is the model's read on the fixture. The amber badge says quoting is switched off while the window is on public display.
The live Lure top bar at 2560 pixels wide: a status ribbon reading uptime, heap, machine RAM, CPU, websocket queue, predictor date, scraped date and the running build commit; then the logo, a rail of nineteen sport filters in two rows, a search field, and two clusters of toolbar controls
The window's own health line, live as this page was built: uptime, heap, machine memory, CPU, the websocket queue, the predictor and scrape dates, and the commit the running build was cut from. Then nineteen sport filters, search and the toolbars. The bar is a single CSS grid — logo | health | centre | toolbar | conductor | power | history — after an earlier version pinned controls to viewport-relative offsets and collided below 2560px.

What it is not

It is not a tipster service. The only claim this page makes about results is the one the bookmakers made for me: around twenty of them banned my accounts. There are no profit, loss, yield or strike-rate figures anywhere on it, and the screenshots were chosen so that none appear. What is on show is the machinery: the models and their features, feed coverage, event reconciliation, the model pipeline and the operations console.

It is not a wrapper around somebody's odds API. Each of the forty-five sources has its own connector in this repository — REST pollers, order-book WebSockets, socket.io clients, HTML scrapers behind rotating proxies, and off-screen browser gateways for the two venues that will not answer a plain HTTP client.

It is not open source and it is not a team product. One person designed it, wrote it, deployed it and gets the pager. The public site at lure.ie is a read-only relay: writes, balances, positions, account connectors, trading configuration and the operations console are all refused before they leave the box.

The models · the heart of it

Nine years of beating the Counter-Strike markets.

The first Counter-Strike models went live in March 2017 and beat the markets for years, until around twenty bookmakers had banned my accounts. Dota 2 and League of Legends followed in 2018, tennis in 2019, Valorant in 2024, StarCraft 2 in 2025 and Mobile Legends in 2026. The edge was never the algorithm. It was the features, and the rule that no feature may use information that did not exist when the match was played.

Feature engineering at depth

The Counter-Strike map model starts from 1,845 candidate features over about 142,000 historical maps: ratings for teams, maps and round handicaps from minus ten to plus ten, per-player ratings on kills, damage, KAST and headshots, HLTV opening duels, trades and clutches, and a large economy block built from round-level scorebot data, buy against buy on each side. Attribute selection cuts it to the 113 that carry signal. Valorant runs the same way on 725.

The veto decides the match

A Counter-Strike series is played on maps the two teams pick and ban. The veto-aware model predicts the likely map sequence from each team's pool, prices every map with its own model, and combines them into the series price, with an endpoint that explains its reasoning map by map.

Measured the only honest way

Validation is chronological: Brier score against a baseline on a later held-out period, and walk-forward backtests, never shuffled folds. Every price the model produced is stored beside what the market was quoting at the same minute, so it can be judged against the market after the fact.

Weka, kept boring

Per-sport runner definers build Weka classifiers — logistic regression under attribute selection and filtering — from replayed historical events. The dependency is pinned rather than floating, because four Weka add-ons declare open version ranges that Maven can never satisfy locally and quietly added fifteen minutes of remote metadata fetches to every build.

weka.classifiers.functions.Logistic · AttributeSelectedClassifier · FilteredClassifier · NumericToNominal

TrueSkill, extended

Team, map and player skill is carried by a Bayesian rating system ported into the codebase from the open-source jskills implementation of TrueSkill and extended to round handicaps and per-player statistics — the factor graph, the Gaussian factors, the layer schedule and the partial-play support, about 3,300 lines across 50 classes, so ratings can be replayed and persisted with the rest of the situational state.

Temporal integrity is the cardinal rule

A player's current rating must never be attached to a historical match. Replay is in true chronological order, which for tournament data means synthesising a round-based day offset because the source stamps every match with the tournament start date. Backtests are strictly online: predict from state built only from earlier matches, then update.

A classifier date is not evidence of freshness

Every published predictor directory carries a training-data watermark: the newest non-future historical event the training replay actually fed it. The nightly refuses to publish a sport whose watermark is older than a configured limit, and a missing stamp fails too. At runtime the API bans trading per event class on a stale watermark, alarms, and paints the freshness chips red in the UI.

That gate exists because of a real failure: a site started serving minified HTML with no whitespace text nodes, a parser that indexed text nodes by position silently stopped importing map results, and the nightly reported success for twenty-seven consecutive nights. The date on the classifier directory was current every one of those mornings.

The pipeline

From forty-five disagreeing books to one event graph.

Every source emits into the same pipeline. Only after normalisation and mapping does anything downstream — pricing, arbitrage, execution, the UI — get to see it.

45 SOURCE CONNECTORS · 37 EMIT PRICES Exchanges · 3 Betfair · Smarkets · Polymarket back and lay, order books Sportsbooks · 23 bet365 · Unibet · Ladbrokes … back-only, two behind browsers Price-only odds Pinnacle observed, never placed on Match data · 8 data-only HLTV · PandaScore · OpenDota Sackmann · Flashscore · rosters Conductor separate JVM · port 8090 01:00 nightly, per sport: · scrape history · gate on the FULL test suite · gate on free disk · retrain + replay · gate on data watermark · publish classifier dir · rebuild jar, restart API a stale watermark bans that sport from trading, not the run Per-sport normalisers one emitter per source · canonical market and runner keys · side orientation · scores and incidents Event mapper · name equality · id counter decides that a row from one book is the same event as a row from another, and stamps the canonical event id unmapped rows are surfaced in the UI rather than silently dropped · World Events maps the same real-world question across venues Canonical event graph Event → Market → Runner → Price one graph per event, all venues inside it Write-behind cache → MySQL + heavy-data files gzipped serialised graphs in a LONG BLOB · HikariCP pool the model classes ARE the persisted format event bus · every update, once Greenlight gate · trader per-event pre-trade check, fail-closed stale-training-data block per sport blocks writes, never inspection Cross-venue arbitrage engine detect → suppress re-fire → re-poll both books → re-price → size → fire FOK a single-venue candidate is rejected Runner predictor Weka logistic classifiers TrueSkill-style ratings + state 8 sports have trained models publishes the predictor directory Account connectors · 8 venues Betfair · Smarkets · Polymarket · Sportsbet.io · Boylesports · Thunderpick · EGB · Pinnacle 3 are executable for arbitrage legs; gateway-backed venues refresh balance off the caller thread Settlement and records every take, close and orphan persisted with venue order ids on submitted legs API · Jetty + Jersey 113 REST endpoints · WebSocket on /websocket lure-ui · desktop 127.0.0.1:8081 · full control lure.ie · public, read-only Caddy relay on EC2 → SSH tunnel → the same API

One emitter per source

Forty-five connector packages, thirty-seven of which emit prices. They are not variations on one HTTP client: an exchange order-book stream, a socket.io market feed, a REST poller, a scraper behind rotating egress proxies and an off-screen browser gateway are all different problems, and each has its own failure mode to surface.

betfair · smarkets · polymarket · pinnacle · bet365 · unibet · ladbrokes · skybet · williamhill · boylesports · sportsbetio · cloudbet · thunderpick · ggbet · csgoempire · hltv · pandascore · opendota · sackmann · flashscore · vlrgg · oracleselixir · leaguepedia · and more

The model classes are the database

Canonical event graphs are Java-serialised and gzipped into a MySQL LONG BLOB, with the heavy parts written out beside it. That makes the data model a persisted wire format: renaming a field, changing a superclass or dropping a serialVersionUID is a migration, not a refactor, and risky changes are proved by deserialising live rows before deploy.

Unmapped is a first-class state

A row that cannot be tied to a canonical event is not discarded. It is held as unmapped, surfaced under its own top-bar filter with a name-equivalency editor next to it, and reconciled when a mapping arrives. Silent dropping is how a platform quietly stops covering a venue.

The Conductor

It retrains and redeploys itself, every night, unattended.

A separate always-on Java process that owns the platform's lifecycle: it scrapes new history, runs the full test suite, retrains per sport, checks the data watermark, publishes the predictor directory, rebuilds the jar and restarts the live API. It runs at 01:00 and nobody watches it.

The Conductor's activity panel: batch daily-20260915-010139, csgo done, published today csgo valorant lol dota2 tennis starcraft2 mobilelegends, and a live training log showing a Weka logistic classifier being built from 17,455 instances and the training-data watermark being written
The operations console this morning. Last night's batch, the sports whose predictors were republished today, and the tail of the training log — a Weka logistic classifier being built from 17,455 replayed instances, then the training-data watermark and situational state persisted to disk.
Conductor nightly schedule controls: update history first, promote and restart when done, earliest days to scrape 7, max training-data age 2 days, prep heap 10G, nightly run enabled at 01:00, and next run 2026-09-16T01:00 with last run 2026-09-15
The schedule, as configured. Maximum training-data age in days is the watermark gate; "promote & restart when done" is the part that makes it a deployment rather than a training job.
Conductor predictor directories: per-sport dropdowns choosing which published classifier directory each sport is served from, with csgo, valorant, lol, dota2, tennis, starcraft2 and mobilelegends each showing seven available directories dated 2026-09-15
Predictor directories. Each sport is served from a dated, published classifier directory, and any of them can be rolled back to an earlier one from here — which is the whole reason training writes to a new directory rather than over the live one.
Conductor feeds panel: Betfair, Polymarket and Smarkets grouped as Exchange, Crypto AI as Prediction book, Pinnacle as Odds, Boylesports and Sportsbet.io as Bookmaker, each with an on switch and a poll interval in seconds
Feed control. Toggles and poll intervals persist outside the jar and are re-injected as JVM arguments on every restart, so a feed switched off at three in the morning is still off after the nightly rebuilds and relaunches the API.

Gates before work, not after

The pipeline refuses to start training if the full Maven test suite fails or if free disk is below its threshold. Both gates are deliberately before the expensive part: there is no point spending hours retraining to fail at publish.

No shell scripts

The lifecycle used to be PowerShell. It is now native Java that shells out only to real tools — git, mvn, jps, netstat, taskkill, java. Process discovery identifies the API by its main jar token, port to PID goes through netstat, and timed-out children are tree-killed through the process handle.

Training outlives its parent

A run is launched as a detached JVM, so restarting or updating the Conductor does not kill a training pass in flight. The scheduler thread is separate from the console's HTTP server, which is how a wedged console once kept firing nightlies for two days while looking completely dead.

Cross-venue arbitrage · a 2026 experiment

Two books, or it is not an arb.

The newest and smallest part of the platform, and an experiment that did not pay its way: in 2026 I tried locking in cross-venue prices instead of betting the models' view. Its safety rules are still worth showing. The engine will not fire both legs on one venue. That rule is enforced in the maths, not in a comment, and it is the single most useful thing in the module.

/*
 * ... That is never a lockable arb: NO venue crosses its own book. A bookmaker prices its own
 * line with a positive overround; an exchange's best asks across complementary outcomes also
 * always rest at or above 100% ... So a single-venue sub-100% "book" is always a stale/laggy
 * snapshot or a data artefact ... and placing both legs on one venue throws away arbitrage's
 * whole point: two INDEPENDENT counterparties, so one side suspending or erroring cannot leave
 * a naked leg. A real arb must span two distinct sources; a single-venue candidate is rejected.
 */
private static boolean isSingleVenue(ArbCandidate candidate) { ... }

From ArbMath.java. The comment is the design note; the method is the gate.

Detect, then distrust

A candidate found on the event thread is not fired from it. The engine suppresses a re-fire until the odds actually move — not just because its own order moved the depth — then, off the hot path, re-polls both books in parallel, re-prices, and sizes to the most volume that keeps every outcome non-negative inside cached balances and real book depth.

Parallel when fresh, sequenced when not

If every leg is fresh on its live feed the legs go in parallel. If a leg had to be re-polled the fire is sequenced hedge-first, so a re-polled leg that misses withholds the taker instead of leaving it naked. Each back leg is submitted at its break-even floor rather than the ideal target, so a price that drifted inside the profitable band still fills.

A one-sided fill is held, not flattened

If only one leg takes, the position is held and flagged ORPHANED — never auto-closed, because closing pays a spread for nothing. Orphans are instead prevented up front: a venue that could not submit right now, because it is disabled or rate-limited, is excluded from leg selection before a candidate is built.

Who may hold a leg

Executable and observation sources are separate lists in configuration, and the header arb flag is restricted to exchanges only. A price-only feed or a back-only sportsbook cannot be hedged out of, so counting one as an arb leg would light a signal the engine would never take.

arb.executableSources=polymarket,betfair,smarkets
arb.observationSources=pinnacle
arbitrage.hedgeablePriceSources=polymarket,betfair,smarkets

Pinnacle is read for price discovery and never placed on. Sportsbet.io and Boylesports have account connectors but reach their venues through browser gateways, so they are not arbitrage legs.

Where the limit actually is

The constraint on cross-venue coverage is not clever maths — it is overlap. An arbitrage needs two independent books quoting the same runner of the same market of the same event at the same moment, which makes the mapping layer, not the engine, the thing that decides how many opportunities exist at all.

That is why the platform spends so much of itself on identity: name equivalencies, per-sport market-key converters, a reverse sweep that reconciles price-only books at mint time, and a World Events layer that maps the same real-world question across venues that would otherwise never meet.

The window at lure.ie

A window onto the engine, not the engine.

A word on what you are looking at when you open lure.ie: it holds no state and places no trades. It is the client I wrote so I could check on the engine from a hotel, and it is switched on now so you can look in. Everything that matters happens in the Java behind it. lure-ui itself is a browser client with no framework: jQuery, Babel and about 37,000 lines of my own JavaScript against the same REST and WebSocket API the desktop uses. Every row has to show sport, source coverage, model state, prices, exposure and controls without opening ten panels.

The live Lure event list on lure.ie: rows of Counter-Strike fixtures for today, each with a sport badge, both team names, a strip of source-coverage and status icons, a start time and coloured source stripes at the right edge. The account column on the left of every row is blank.
What the window actually shows, captured this morning through lure.ie. One row per canonical event, each carrying its sport, its source coverage, its model state and its start time. Row colour is the model's read on the fixture, not a result. The left column of every row is blank on purpose: that is where exposure sits on my own screen, and the relay refuses position data before it leaves the house, so the window shows you the market and never the book I am running against it.
The Match Winner market of baks versus spirit academy opened on lure.ie: baks priced at 1.28 with the model at 78 percent, spirit academy at 5 with the model at 20 percent. Under each runner a chart on a shared probability axis shows the resting depth at every price as teal and pink bars for polymarket, pinnacle and betfair, and the model's prediction as a dotted bell curve to the right of the market for baks and to the left of it for spirit academy
The Match Winner market opened, with the model in the picture. Each runner header carries the best price and the model's probability: baks at 1.28 against a 78% read, spirit academy at 5.00 against 20%. Under each one, on a single probability axis, the solid bars are the resting depth at every price on polymarket, pinnacle and betfair, and the dotted curve is the model's prediction with its uncertainty. Here the model sits to the right of the money on baks and to the left of it on spirit academy, which is the whole conversation this screen exists to have. Captured through lure.ie itself.
The ARB WebSockets panel: counters for tracked, pre-live, covered, hot and average score; a table of betfair, pinnacle and polymarket with role exec or observe, tracked ids, covered, hot and stream score; then a list of markets with a score, event, sport, market, start time, book percentage, recent move, how many sources are mapped and priced, and a state of hot or mid covered
Cross-venue coverage, live this afternoon. 6,417 market streams tracked, 53 currently covered on both sides, 28 hot. Each row is a market the engine is watching across venues with its book percentage, how it has moved, and how many sources are mapped and priced against how many can actually be executed on. betfair and polymarket carry role exec; pinnacle is observe and is never placed on.
The Pipeline Metrics panel: trader queue wait, batch process and queue depth tiles, an events-published counter, a startup and warmup timeline listing first cycle and first publish for betfair, flashscoreLive, vlr.gg, pinnacle, smarkets, pandascore, gosugamers-mobilelegends, Leaguepedia, hltv.org and polymarket, and a per-source emitter cycle table
Where the delays are. The warm-up timeline stamps every feed's first cycle and first publish against process start — betfair at 41 seconds, pinnacle at 1m07s, hltv.org at 7m02s — and the table below it tracks each emitter's cycle time since. This is the panel I open first when the engine feels slow, and the one that tells me which of forty-five sources is the reason.
The historic view of Spirit versus MOUZ from BLAST Open Porto on 6 September, Match Winner market: for each runner a time-series chart from 3 PM to 6 PM plotting the polymarket, betfair, pinnacle and smarkets prices and the model's prediction as stepped lines on a probability axis, with a dense row of match-incident markers along the baseline for kills, bomb plants and round ends
The runner chart, in the historic view. Spirit against MOUZ at BLAST Open Porto on 6 September, Match Winner, the whole match from 15:00 to 18:00. The stepped lines are what polymarket, betfair, pinnacle and smarkets were quoting minute by minute, on the same probability axis as the model's line, and the markers along the baseline are the scorebot feed: every kill, bomb plant and round end the engine recorded while it was pricing. Spirit's price walks from the low sixties into the eighties as the maps go by; MOUZ's mirrors it. Every point on this chart is a stored, timestamped observation, which is what makes the model measurable against the market after the fact.
The same match, Map 1 Winner market in the historic view: each runner's chart runs from the start of the map at 3 PM until the map ends just after 4 PM, with the price line moving as rounds are won and the incident markers stopping at the flag that marks the end of the map
The same match one level down: Map 1 Winner. The map's own market lives for an hour, the price moves round by round with the incidents underneath it, and both stop at the flag that marks the end of the map. Map-scoped incidents attach only to map-scoped charts; the match chart above carries all three maps.
The Polymarket WebSocket panel pinned open from the WS chip on lure.ie: connected, 225 connected events, 1,352 of 1,600 assets, 8 sockets of 7 wanted, then a table of subscribed events with sport, name, date, last message time, last price change and a sixty-second sparkline of messages arriving for each
The market socket, as the messages arrive. The WS chip in the top bar opens this panel: one exchange socket per two hundred assets, 225 events subscribed, and for each one the second-by-second bars of messages coming in over the last minute, with the time of the last message and the last actual price change beside them. This is the feed the depth charts above are drawn from, and the panel I look at when a price seems stale.

Why there is no framework

It started in August 2016 and has been rewritten in place ever since. A trading screen that has to paint hundreds of live rows from a WebSocket, keep per-row subscriptions open, and survive a reconnect without losing the user's filters is mostly a state problem, and the framework of 2016 would have been the framework of 2016.

Market and runner labels never surface raw canonical keys; twelve per-sport key converters turn them into the words a trader uses. Source, provider and platform names all go through one chip builder, so a venue looks the same everywhere it appears.

How lure.ie is public without being open

The public site is served by Caddy on a small EC2 box, which reaches the home server's loopback API through an outbound SSH tunnel. The relay holds an allow-list: writes, credit balances, positions, account connectors, trading configuration, proxy credentials and every operations endpoint are refused at the relay, and the Conductor is not relayed at all. Ask it for market detail and it answers the public view is read-only; ask it for arbitrage history and it answers not available on the public view. That is the design, not an outage.

It is honest about its own limits, too. The tunnel is started from the Windows startup folder, so the public site is live only while the machine is logged in — which is exactly the sort of thing a page like this should say out loud.

Engineering

Ten years, one pair of hands.

Backend222,000 lines of Java across 922 files, one Maven module, on Spring, Jetty, Jersey and Hibernate
Front end37,000 lines of my own JavaScript across 25 files, plus 12,000 lines of CSS; jQuery and Babel, no framework
Commits5,718 on the backend since 1 April 2016, 1,052 on the front end since 1 August 2016. Thirteen of the 6,770 carry an author name that is not one of mine
Feeds45 source connectors; 37 emit prices, 8 are match-data only, 8 carry an account connector
Sports20 served in the live API; 8 have trained classifiers; 12 key converters in the UI
API113 REST endpoints and a WebSocket, on 8081 plain and 8043 TLS
Arbitrage engine8,800 lines across 21 classes: detection maths, execution, settlement, records, exposure
Conductor7,100 lines across 18 classes: lifecycle, training pipeline, builder, cleanup, worktrees
Tests468 JUnit tests across 61 classes; the nightly refuses to train if any of them fail
StorageMySQL with a write-behind cache in front of it, serialised event graphs in blobs, heavy data on the filesystem, retention and stale-row reapers
OperationsOne Windows host in Cork, a scheduled logon task, a small EC2 box for the public relay and the tailnet control plane

Where the numbers come from

Line and file counts are git ls-files over tracked sources, counted on 15 September 2026; the front-end figure excludes the bundled jQuery, jQuery UI and Tooltipster that ship in the same tree. Commit counts and first-commit dates are from git log. Endpoint, test, feed and class counts are grepped from the code. The venue lists are the live configuration values, not a wish list.

Prior art, paid

I did this professionally before I did it alone. At Gravity (2012–2018), the trading business in the Matchbook group that later became Newton Squared 90, I designed, built and owned the market-making and risk platform that priced every event on the Matchbook exchange and traded multi-million-dollar accounts across many sports. At RISQ Capital (2018–2019) I worked on an ultra-low-latency tennis trading system in C++. At Pythia Sports (to April 2023) I worked on sports trading systems remotely. Lure is what happens when the same person gets to make every call.

Timeline

  1. Gravity, later Newton Squared 90, Cork. The trading side of the Matchbook group, and the market-making and risk platform I designed and owned.

  2. First commit on this repository, while still at Gravity. It has never stopped.

  3. lure-ui begins: the browser client that is still the front end today.

  4. The first Counter-Strike models go live. They beat the markets for years, until around twenty bookmakers had banned my accounts.

  5. Dota 2 and League of Legends models in 2018, tennis in 2019, Valorant in 2024, StarCraft 2 in 2025, Mobile Legends in 2026.

  6. RISQ Capital, London. Ultra-low-latency tennis trading in C++.

  7. Pythia Sports, remote. Left salaried work in April 2023 and have been independent since.

  8. The Conductor replacing Jenkins outright, the World Events mapper, the public read-only relay at lure.ie, and a cross-venue arbitrage experiment that did not pay its way.

Honest limits

One person, no second pair of eyes, and no code review that is not my own. The primary host is Windows and nothing else is a supported target. The public relay is live only while that machine is logged in. Model coverage is uneven — eight sports have trained classifiers and the rest are priced from the market. And a platform this old carries its history: parts of it are 2016 code that works and has not been worth rewriting.

Contact

Ken Barry

Fifteen years of Java on real-time trading systems, ten of them on this one. Available immediately, remote, based in Cork.