Taps can be unreliable inside this app's built-in browser. For the full experience, use its menu to open this page in your browser.
Skip to content

A Guard That Blocked Its Own AI

A devlog by Neeraj Kotwani, who builds every game on Elipar.

Neon Four is a four-in-a-row drop game with hot-seat two-player and a vs-CPU mode. The engine is small and the AI is the interesting part — right up until I shipped a one-line guard that made the single-player mode impossible to play, and watched every test pass anyway.

The line

Dropping a disc goes through one function. It had this at the top:

if (mode === 'cpu' && currentPlayer === 2) return false;

The intent is obvious and reasonable: in vs-CPU mode, a human tap during the CPU's turn should do nothing.

The problem is that this is the same function the AI uses to place its own disc. The AI's move is scheduled on a timer so the opponent reads as thinking rather than snapping back instantly — and by the time that timer fires, the current player is player 2. Which is the exact condition the guard rejects.

So the CPU's own legitimate move was refused. Every vs-CPU game froze forever after the human's first move, showing "CPU THINKING...", with the CPU never placing anything. Half the modes in the game, completely dead.

Every isolated test passed, and that is the real story

I had tests. Win detection in all four directions. The AI takes an available winning move. The AI blocks a real single threat. Search depth timings at five different depths. All green, all meaningless — because not one of them drove a real game past a single CPU turn.

They each called the engine directly with a constructed board. The broken path only exists in the seam between the input handler, the timer and the drop function, and nothing was exercising that seam.

What found it was a deterministic full-game test: play to completion, forcing the AI's timer and the drop animation forward instead of waiting in real time, and assert the game reaches a terminal state. It found the move count stuck at 1 after a hundred loop iterations.

That is the second time on this site that an end-to-end test has caught something the unit tests actively gave false confidence about. The curling game had a sibling bug — a CPU turn that re-entered its own scheduling branch on every animation frame and queued about thirty extra throws — and its isolated tests were all green too. If a game has a CPU opponent, one test must play a whole game.

The fix is where the guard lives, not whether it exists

The guard is correct. It is in the wrong place.

It is now in the pointerdown handler, where it suppresses human input during the AI's turn and has no opinion about the AI calling the drop function directly. The shared game-logic function went back to being about game logic.

That is a pattern worth naming, because it keeps recurring: a check that is really about the user interface drifts down into the shared function, and the next caller that is not a user trips over it. The hover-preview branch in the same file carries the same condition, for the same reason, and that one is right where it is — it genuinely is a UI concern.

There is a comment on the handler explaining this, pointing at the failure it caused. Not for me — so that the next person who notices the guard looks oddly placed does not helpfully move it back.

Search depth 10, because I timed it

The opponent is minimax with alpha-beta pruning: score the board on centre-column control plus a windowed heuristic over every possible four-in-a-row, order candidate moves centre-outward as a cheap pruning win, and search N moves ahead.

N was not guessed. Timed on an empty board, which is the worst case for branching factor:

  • depth 7 — 19ms
  • depth 8 — 23ms
  • depth 9 — 228ms
  • depth 10 — 289ms
  • depth 11 — 1557ms

Depth 10. It sits comfortably under the 450ms thinking delay that is already there, so the pause costs nothing real, and the empty board is genuinely the worst case — every later move has strictly fewer columns left to search. Depth 11 is four times slower for one more ply and would have been felt.

Measuring that is ten minutes of work and the alternative is picking a number that is either needlessly weak or occasionally janky, with no way to know which.

A second failure, purely my test's fault

The score-reporting test read its results array in the same call that triggered the win, and got either an empty array or a leftover report from the previous scenario.

postMessage is always delivered asynchronously, even to the same window. A synchronous read immediately after triggering it races the browser's own message queue, and which one wins tells you nothing.

Fixed by waiting for delivery, and by clearing the array before each scenario rather than trusting it to be empty from setup. The game was reporting correctly the whole time — exactly one report per game.

I have since hit the identical trap in another game's test suite. It is worth internalising: if a test triggers something that ends in a postMessage, it has to wait.

What it scores

A vs-CPU win pays more the faster you do it, with a floor so a long grind still counts; a draw pays a little; a loss pays nothing. Hot-seat two-player reports no score at all, because there is no single identifiable player on the device — the same rule every other two-player game on this site follows.

Neon Four is free in a browser, two players on one phone or one against the CPU, no account, no download. The CPU takes its turn.

We use cookies to keep games running and, once ads are live, to show relevant content. See our Privacy Policy for details.