Ken Barry · Cork · sports trading systems since 2012 · this repository since April 2016
Lure is the e-sports prediction platform I have built and run alone since 2016. Its heart is machine learning: Counter-Strike models, in use since 2017, that beat the bookmakers for years until around twenty of them had banned my accounts, and now models for Dota 2, League of Legends, Valorant, StarCraft 2, Mobile Legends and tennis. It pulls match data and prices from forty-five sources, reconciles them into one event graph, and retrains and redeploys every model every night without a human. lure.ie is not the product and does no trading: it is the read-only window I built so I could check on the engine while travelling, and it is switched on so you can look in.
What it is
Forty-five venues and data providers each have their own idea of what a match is called, when it starts, which market is which and which side is which. Before any model can be measured against the market, Lure has to decide that two rows from two different books are the same runner in the same market of the same event, and be right every time.
Everything downstream depends on that. A model is only being measured against the market if the runner it priced is the runner the bookmaker is quoting, and a bet is only the bet the model meant if the mapping is correct. Most of the code in this repository exists to earn that identity and then keep it consistent while forty-five sources disagree in real time.
It runs continuously on one Windows machine in Cork, backed by MySQL, with a read-only monitoring window relayed out to lure.ie. I designed, wrote, deployed and operate all of it. Across 6,770 commits on the two repositories, thirteen carry anyone's name but mine — recent ones stamped by a coding assistant working on my machine.
logo | health | centre | toolbar | conductor | power | history — after an earlier version pinned controls to viewport-relative offsets and collided below 2560px.It is not a tipster service. The only claim this page makes about results is the one the bookmakers made for me: around twenty of them banned my accounts. There are no profit, loss, yield or strike-rate figures anywhere on it, and the screenshots were chosen so that none appear. What is on show is the machinery: the models and their features, feed coverage, event reconciliation, the model pipeline and the operations console.
It is not a wrapper around somebody's odds API. Each of the forty-five sources has its own connector in this repository — REST pollers, order-book WebSockets, socket.io clients, HTML scrapers behind rotating proxies, and off-screen browser gateways for the two venues that will not answer a plain HTTP client.
It is not open source and it is not a team product. One person designed it, wrote it, deployed it and gets the pager. The public site at lure.ie is a read-only relay: writes, balances, positions, account connectors, trading configuration and the operations console are all refused before they leave the box.
The models · the heart of it
The first Counter-Strike models went live in March 2017 and beat the markets for years, until around twenty bookmakers had banned my accounts. Dota 2 and League of Legends followed in 2018, tennis in 2019, Valorant in 2024, StarCraft 2 in 2025 and Mobile Legends in 2026. The edge was never the algorithm. It was the features, and the rule that no feature may use information that did not exist when the match was played.
The Counter-Strike map model starts from 1,845 candidate features over about 142,000 historical maps: ratings for teams, maps and round handicaps from minus ten to plus ten, per-player ratings on kills, damage, KAST and headshots, HLTV opening duels, trades and clutches, and a large economy block built from round-level scorebot data, buy against buy on each side. Attribute selection cuts it to the 113 that carry signal. Valorant runs the same way on 725.
A Counter-Strike series is played on maps the two teams pick and ban. The veto-aware model predicts the likely map sequence from each team's pool, prices every map with its own model, and combines them into the series price, with an endpoint that explains its reasoning map by map.
Validation is chronological: Brier score against a baseline on a later held-out period, and walk-forward backtests, never shuffled folds. Every price the model produced is stored beside what the market was quoting at the same minute, so it can be judged against the market after the fact.
Per-sport runner definers build Weka classifiers — logistic regression under attribute selection and filtering — from replayed historical events. The dependency is pinned rather than floating, because four Weka add-ons declare open version ranges that Maven can never satisfy locally and quietly added fifteen minutes of remote metadata fetches to every build.
Team, map and player skill is carried by a Bayesian rating system ported into the codebase from the open-source jskills implementation of TrueSkill and extended to round handicaps and per-player statistics — the factor graph, the Gaussian factors, the layer schedule and the partial-play support, about 3,300 lines across 50 classes, so ratings can be replayed and persisted with the rest of the situational state.
A player's current rating must never be attached to a historical match. Replay is in true chronological order, which for tournament data means synthesising a round-based day offset because the source stamps every match with the tournament start date. Backtests are strictly online: predict from state built only from earlier matches, then update.
Every published predictor directory carries a training-data watermark: the newest non-future historical event the training replay actually fed it. The nightly refuses to publish a sport whose watermark is older than a configured limit, and a missing stamp fails too. At runtime the API bans trading per event class on a stale watermark, alarms, and paints the freshness chips red in the UI.
That gate exists because of a real failure: a site started serving minified HTML with no whitespace text nodes, a parser that indexed text nodes by position silently stopped importing map results, and the nightly reported success for twenty-seven consecutive nights. The date on the classifier directory was current every one of those mornings.
The pipeline
Every source emits into the same pipeline. Only after normalisation and mapping does anything downstream — pricing, arbitrage, execution, the UI — get to see it.
Forty-five connector packages, thirty-seven of which emit prices. They are not variations on one HTTP client: an exchange order-book stream, a socket.io market feed, a REST poller, a scraper behind rotating egress proxies and an off-screen browser gateway are all different problems, and each has its own failure mode to surface.
Canonical event graphs are Java-serialised and gzipped into a MySQL LONG BLOB, with the heavy parts written out beside it. That makes the data model a persisted wire format: renaming a field, changing a superclass or dropping a serialVersionUID is a migration, not a refactor, and risky changes are proved by deserialising live rows before deploy.
A row that cannot be tied to a canonical event is not discarded. It is held as unmapped, surfaced under its own top-bar filter with a name-equivalency editor next to it, and reconciled when a mapping arrives. Silent dropping is how a platform quietly stops covering a venue.
The Conductor
A separate always-on Java process that owns the platform's lifecycle: it scrapes new history, runs the full test suite, retrains per sport, checks the data watermark, publishes the predictor directory, rebuilds the jar and restarts the live API. It runs at 01:00 and nobody watches it.
The pipeline refuses to start training if the full Maven test suite fails or if free disk is below its threshold. Both gates are deliberately before the expensive part: there is no point spending hours retraining to fail at publish.
The lifecycle used to be PowerShell. It is now native Java that shells out only to real tools — git, mvn, jps, netstat, taskkill, java. Process discovery identifies the API by its main jar token, port to PID goes through netstat, and timed-out children are tree-killed through the process handle.
A run is launched as a detached JVM, so restarting or updating the Conductor does not kill a training pass in flight. The scheduler thread is separate from the console's HTTP server, which is how a wedged console once kept firing nightlies for two days while looking completely dead.
Cross-venue arbitrage · a 2026 experiment
The newest and smallest part of the platform, and an experiment that did not pay its way: in 2026 I tried locking in cross-venue prices instead of betting the models' view. Its safety rules are still worth showing. The engine will not fire both legs on one venue. That rule is enforced in the maths, not in a comment, and it is the single most useful thing in the module.
/*
* ... That is never a lockable arb: NO venue crosses its own book. A bookmaker prices its own
* line with a positive overround; an exchange's best asks across complementary outcomes also
* always rest at or above 100% ... So a single-venue sub-100% "book" is always a stale/laggy
* snapshot or a data artefact ... and placing both legs on one venue throws away arbitrage's
* whole point: two INDEPENDENT counterparties, so one side suspending or erroring cannot leave
* a naked leg. A real arb must span two distinct sources; a single-venue candidate is rejected.
*/
private static boolean isSingleVenue(ArbCandidate candidate) { ... }
From ArbMath.java. The comment is the design note; the method is the gate.
A candidate found on the event thread is not fired from it. The engine suppresses a re-fire until the odds actually move — not just because its own order moved the depth — then, off the hot path, re-polls both books in parallel, re-prices, and sizes to the most volume that keeps every outcome non-negative inside cached balances and real book depth.
If every leg is fresh on its live feed the legs go in parallel. If a leg had to be re-polled the fire is sequenced hedge-first, so a re-polled leg that misses withholds the taker instead of leaving it naked. Each back leg is submitted at its break-even floor rather than the ideal target, so a price that drifted inside the profitable band still fills.
If only one leg takes, the position is held and flagged ORPHANED — never auto-closed, because closing pays a spread for nothing. Orphans are instead prevented up front: a venue that could not submit right now, because it is disabled or rate-limited, is excluded from leg selection before a candidate is built.
Executable and observation sources are separate lists in configuration, and the header arb flag is restricted to exchanges only. A price-only feed or a back-only sportsbook cannot be hedged out of, so counting one as an arb leg would light a signal the engine would never take.
arb.executableSources=polymarket,betfair,smarkets
arb.observationSources=pinnacle
arbitrage.hedgeablePriceSources=polymarket,betfair,smarkets
Pinnacle is read for price discovery and never placed on. Sportsbet.io and Boylesports have account connectors but reach their venues through browser gateways, so they are not arbitrage legs.
The constraint on cross-venue coverage is not clever maths — it is overlap. An arbitrage needs two independent books quoting the same runner of the same market of the same event at the same moment, which makes the mapping layer, not the engine, the thing that decides how many opportunities exist at all.
That is why the platform spends so much of itself on identity: name equivalencies, per-sport market-key converters, a reverse sweep that reconciles price-only books at mint time, and a World Events layer that maps the same real-world question across venues that would otherwise never meet.
The window at lure.ie
A word on what you are looking at when you open lure.ie: it holds no state and places no trades. It is the client I wrote so I could check on the engine from a hotel, and it is switched on now so you can look in. Everything that matters happens in the Java behind it. lure-ui itself is a browser client with no framework: jQuery, Babel and about 37,000 lines of my own JavaScript against the same REST and WebSocket API the desktop uses. Every row has to show sport, source coverage, model state, prices, exposure and controls without opening ten panels.
exec; pinnacle is observe and is never placed on.It started in August 2016 and has been rewritten in place ever since. A trading screen that has to paint hundreds of live rows from a WebSocket, keep per-row subscriptions open, and survive a reconnect without losing the user's filters is mostly a state problem, and the framework of 2016 would have been the framework of 2016.
Market and runner labels never surface raw canonical keys; twelve per-sport key converters turn them into the words a trader uses. Source, provider and platform names all go through one chip builder, so a venue looks the same everywhere it appears.
The public site is served by Caddy on a small EC2 box, which reaches the home server's loopback API through an outbound SSH tunnel. The relay holds an allow-list: writes, credit balances, positions, account connectors, trading configuration, proxy credentials and every operations endpoint are refused at the relay, and the Conductor is not relayed at all. Ask it for market detail and it answers the public view is read-only; ask it for arbitrage history and it answers not available on the public view. That is the design, not an outage.
It is honest about its own limits, too. The tunnel is started from the Windows startup folder, so the public site is live only while the machine is logged in — which is exactly the sort of thing a page like this should say out loud.
Engineering
| Backend | 222,000 lines of Java across 922 files, one Maven module, on Spring, Jetty, Jersey and Hibernate |
|---|---|
| Front end | 37,000 lines of my own JavaScript across 25 files, plus 12,000 lines of CSS; jQuery and Babel, no framework |
| Commits | 5,718 on the backend since 1 April 2016, 1,052 on the front end since 1 August 2016. Thirteen of the 6,770 carry an author name that is not one of mine |
| Feeds | 45 source connectors; 37 emit prices, 8 are match-data only, 8 carry an account connector |
| Sports | 20 served in the live API; 8 have trained classifiers; 12 key converters in the UI |
| API | 113 REST endpoints and a WebSocket, on 8081 plain and 8043 TLS |
| Arbitrage engine | 8,800 lines across 21 classes: detection maths, execution, settlement, records, exposure |
| Conductor | 7,100 lines across 18 classes: lifecycle, training pipeline, builder, cleanup, worktrees |
| Tests | 468 JUnit tests across 61 classes; the nightly refuses to train if any of them fail |
| Storage | MySQL with a write-behind cache in front of it, serialised event graphs in blobs, heavy data on the filesystem, retention and stale-row reapers |
| Operations | One Windows host in Cork, a scheduled logon task, a small EC2 box for the public relay and the tailnet control plane |
Line and file counts are git ls-files over tracked sources, counted on 15 September 2026; the front-end figure excludes the bundled jQuery, jQuery UI and Tooltipster that ship in the same tree. Commit counts and first-commit dates are from git log. Endpoint, test, feed and class counts are grepped from the code. The venue lists are the live configuration values, not a wish list.
I did this professionally before I did it alone. At Gravity (2012–2018), the trading business in the Matchbook group that later became Newton Squared 90, I designed, built and owned the market-making and risk platform that priced every event on the Matchbook exchange and traded multi-million-dollar accounts across many sports. At RISQ Capital (2018–2019) I worked on an ultra-low-latency tennis trading system in C++. At Pythia Sports (to April 2023) I worked on sports trading systems remotely. Lure is what happens when the same person gets to make every call.
Gravity, later Newton Squared 90, Cork. The trading side of the Matchbook group, and the market-making and risk platform I designed and owned.
First commit on this repository, while still at Gravity. It has never stopped.
lure-ui begins: the browser client that is still the front end today.
The first Counter-Strike models go live. They beat the markets for years, until around twenty bookmakers had banned my accounts.
Dota 2 and League of Legends models in 2018, tennis in 2019, Valorant in 2024, StarCraft 2 in 2025, Mobile Legends in 2026.
RISQ Capital, London. Ultra-low-latency tennis trading in C++.
Pythia Sports, remote. Left salaried work in April 2023 and have been independent since.
The Conductor replacing Jenkins outright, the World Events mapper, the public read-only relay at lure.ie, and a cross-venue arbitrage experiment that did not pay its way.
One person, no second pair of eyes, and no code review that is not my own. The primary host is Windows and nothing else is a supported target. The public relay is live only while that machine is logged in. Model coverage is uneven — eight sports have trained classifiers and the rest are priced from the market. And a platform this old carries its history: parts of it are 2016 code that works and has not been worth rewriting.
Contact
Fifteen years of Java on real-time trading systems, ten of them on this one. Available immediately, remote, based in Cork.