Benchmark

Traditional vs Jev

The same problems, the same simulation code, the same legal moves. The traditional column is the textbook solution for each problem; the Jev column is one Choice call per decision. Two LLM columns show GPT-5.6 Luna with reasoning off (the like-for-like comparison) and on (a reference); both get exactly Jev's input and must answer with one legal option.

TraditionalJevLuna, reasoning offLuna, reasoning on
Model
jev-1.13.0
Jev calls
10164
Median latency
367 ms
Episodes per setting
10
Runs
19 Sept, 09:50 UTC · ef02480 · eightrank, eightpuzzle, missionaries, vacuum, restaurant, wumpus, tictactoe, pos
Reproduce with npm run benchmark. Counts show a 95% interval underneath; where two intervals overlap widely, treat the difference as noise. Jev is not perfectly deterministic (see Consistency).
ProblemMeasureTraditionalJevLuna, reasoning offLuna, reasoning on
01VacuumCleanliness over 24 stepsin scope
mean of 10 seeded dirt schedules
87%88%88%88%
Moves (energy)in scope15.73.79.115.5
02RestaurantAgrees with AIMA treein scope
40 seeded random scenarios; no ground truth exists
reference30 / 406086%31 / 406288%33 / 406891%
03WumpusEscaped with goldin scope
AIMA world + 49 seeded worlds
29 / 504471%23 / 503360%16 / 502146%23 / 503360%
Mean scorein scope512.1245.5292.6228.4
04MissionariesSolved, current state onlystress test
30-step limit
BFS: 11 moves0 / 10028%0 / 10028%8 / 104994%
Solved, with memoryin scope10 / 1072100%10 / 1072100%10 / 1072100%
Optimal, with lookaheadin scope10 / 1072100%10 / 1072100%10 / 1072100%
05Tic-Tac-ToeNot lost vs perfect play, board onlystress test
plays X and O
minimax: always4 / 20842%7 / 201857%18 / 207097%
Not lost vs perfect play, + tacticsin scope14 / 204885%16 / 205892%19 / 207699%
Won vs random, board onlystress test12 / 203978%13 / 204382%16 / 205892%
07POS TaggingTagging accuracy vs human annotationin scope
EWT test sample, 49 Penn tags
HMM 92.3% · MFT 88.4%92.9%91.794.0%93.6%92.594.6%94.5%93.595.4%
068-PuzzleRanking + search in code: solvedin scope
30 puzzles (depths 8, 12, 16), widening beam 1→13, ≤300 rankings; Manhattan ranker in the classic column
29 / 308399%16 / 303670%12 / 302558%
068-PuzzleA* nodes generated, depth 12stress test
mean of 4 puzzles; Manhattan in the classic column
4369119
Estimates above truthstress test
board + features; above = breaks A* optimality
0% (Manhattan)19%5%
Jev vs GPT-5.6 Luna
  • Luna with reasoning off (10042 calls): median 1026 ms and $0.00014 per call, about 2.8× slower and 3.7× the cost of Jev (367 ms, $0.000038).
  • Luna with reasoning on (1868 calls): median 1563 ms and $0.00062 per call, about 4.3× slower and 16.2× the cost of Jev (367 ms, $0.000038).
  • Planning without help from code (stress tests): Missionaries from the current state alone, solved 0/10 by Jev, 0/10 by Luna with reasoning off and 8/10 with reasoning on; Tic-Tac-Toe from the board alone against perfect play, not lost 4 / 20, 7 / 20 and 18 / 20. The advantage comes with reasoning; with memory or tactics from code the gap mostly closes.
  • On its headline use case, many-label classification (POS tagging, in scope), accuracy is close: Jev 92.9%, Luna with reasoning off 93.6%, with reasoning on 94.5%, against 92.3% for the HMM tagger. The same tagging cost $0.1419 with Jev and $0.5479 with Luna, reasoning off. Jev − Luna off: -0.7 points (95% interval -1.5 to 0.1), not a clear difference. Jev − Luna on: -1.5 points (95% interval -2.4 to -0.8). Jev − the HMM: +0.6 points (95% interval -0.7 to 1.9), not a clear difference.

Same payloads, same legal options, small samples. Luna returns a choice without probabilities. “Reasoning off” is the like-for-like comparison with Jev, which does not reason; “reasoning on” is a reference. With reasoning on, a single 8-Puzzle call took more than three minutes, so the 8-Puzzle was run with reasoning off only.

Economics

Cost and speed

Every call records its tokens and latency. Cost is computed at list price: Jev $0.042 per 1M input tokens (Vercel AI Gateway listing; output not priced); Luna $0.20 input / $1.20 output per 1M (reasoning tokens are billed as output).
Jev
Model
jev-1.13.0
Calls
10164
Median latency
367 ms
Tokens in / out
9,256,585 / 1,521,410
Total cost
$0.3888
Per call
$0.000038
problems: eightrank, eightpuzzle, missionaries, vacuum, restaurant, wumpus, tictactoe, pos
Luna, reasoning off
Model
gpt-5.6-luna
Calls
10042
Median latency
1026 ms
Tokens in / out
6,155,997 / 157,943
Total cost
$1.42
Per call
$0.00014
problems: tictactoe, eightpuzzle, missionaries, eightrank, vacuum, restaurant, wumpus, pos
Luna, reasoning on
Model
gpt-5.6-luna
Calls
1868
Median latency
1563 ms
Tokens in / out
3,356,103 / 407,287
Total cost
$1.16
Per call
$0.00062
problems: tictactoe, vacuum, restaurant, wumpus, pos, missionaries
Fairness

What each solver is given

A benchmark is only as fair as its inputs. For every problem: what the traditional solution uses, and the exact payload Jev receives. The LLM receives the same payload verbatim.
01 · Vacuum
Traditional

The current room's status and location. Two if-statements: dirty → clean, else move.

Jev

Its location and the status of both rooms. More than the rule sees, which is why it can choose to wait.

Example payload, verbatim
{
  state: {
    agentLocation: "A",
    roomA: "dirty",
    roomB: "clean"
  },
  question: "You control a vacuum-cleaning robot in a world of two rooms, A (left) and B (right). Dirt can appear in either room at any time. Which action should the robot take next?",
  actions: ["CLEAN", "MOVE_RIGHT", "WAIT"],
  descriptions: {
    CLEAN: "Vacuum the room the robot is currently in.",
    MOVE_LEFT: "Move left, from room B into room A.",
    MOVE_RIGHT: "Move right, from room A into room B.",
    WAIT: "Stay put and do nothing this step."
  }
}
02 · Restaurant
Traditional

Eight attributes: wait bucket, hunger, alternative, reservation, bar, Fri/Sat, raining (and patrons, fixed at Full). No kids, no ratings.

Jev

Ten plain-language fields including kids and both ratings. Controls on the exhibit also ask with only the tree's attributes.

Example payload, verbatim
{
  state: {
    quotedWaitMinutes: 45,
    hunger: "very hungry",
    raining: false,
    withYoungKids: true,
    hasReservation: false,
    barToWaitAt: false,
    fridayOrSaturdayNight: true,
    thisRestaurantRating: 4.5,
    alternativeRestaurantNearby: true,
    alternativeRestaurantRating: 4.1
  },
  question: "You are advising a party that has just arrived at a full restaurant and been quoted a wait for a table. Given their situation, should they wait for this table or leave?",
  actions: ["WAIT", "LEAVE"],
  descriptions: {
    WAIT: "Stay and wait for a table at this restaurant.",
    LEAVE: "Leave now and eat somewhere else."
  }
}
03 · Wumpus
context not equal
Traditional

The same exact risk estimates as Jev, plus the whole known map: it plans multi-step paths through visited cells to the nearest safe unexplored cell.

Jev

Its own cell, percepts and its neighbouring cells' risks (two to four, depending on position). No map beyond its neighbours. The classic agent has more context here.

Example payload, verbatim
{
  state: {
    position: {
      x: 1,
      y: 0
    },
    entrance: {
      x: 0,
      y: 0
    },
    percepts: {
      breeze: true,
      stench: false,
      glitter: false,
      heardScream: false
    },
    hasArrow: true,
    hasGold: false,
    wumpusKilled: false,
    stepsTaken: 1,
    stepLimit: 40,
    timesVisitedThisCell: 1,
    north: {
      visited: false,
      timesVisited: 0,
      pitRisk: 0.56,
      wumpusRisk: 0,
      knownSafe: false
    },
    east: {
      visited: false,
      timesVisited: 0,
      pitRisk: 0.56,
      wumpusRisk: 0,
      knownSafe: false
    },
    west: {
      visited: true,
      timesVisited: 1,
      pitRisk: 0,
      wumpusRisk: 0,
      knownSafe: true
    }
  },
  question: "You are an agent exploring a dark 4×4 cave to find gold and get out alive. Some cells have bottomless pits or a Wumpus. Given your percepts and the estimated risks of neighbouring cells, which action should you take next?",
  actions: [
    "MOVE_NORTH",
    "MOVE_EAST",
    "MOVE_WEST",
    "SHOOT_NORTH",
    "SHOOT_EAST",
    "SHOOT_WEST",
    "RETREAT"
  ]
}
04 · Missionaries
Traditional

The full transition model. BFS generates and remembers every state until it reaches the goal.

Jev

Depends on the context level: from bank counts and legal moves (with the state each leads to) up to memory, lookahead and distance to goal.

Example payload, verbatim
{
  state: {
    leftBank: {
      missionaries: 2,
      cannibals: 2
    },
    rightBank: {
      missionaries: 1,
      cannibals: 1
    },
    boatIsOn: "right bank",
    goal: "all 3 missionaries and all 3 cannibals on the right bank"
  },
  question: "Three missionaries and three cannibals must all cross a river from the left bank to the right bank. The boat carries one or two people and cannot cross empty. On either bank, if any missionaries are present, cannibals must never outnumber them. Here is the current state and the legal moves available right now. Which move should be made next?",
  actions: ["MOVE_1_MISSIONARY", "MOVE_1_MISSIONARY_1_CANNIBAL"],
  descriptions: {
    MOVE_1_MISSIONARY: "Take 1 missionary from the right bank to the left bank. Afterwards: left bank 3 missionaries and 2 cannibals; right bank 0 missionaries and 1 cannibals; boat on the left bank.",
    MOVE_1_MISSIONARY_1_CANNIBAL: "Take 1 missionary + 1 cannibal from the right bank to the left bank. Afterwards: left bank 3 missionaries and 3 cannibals; right bank 0 missionaries and 0 cannibals; boat on the left bank."
  }
}
Example payload · + memory level
{
  state: {
    leftBank: {
      missionaries: 2,
      cannibals: 2
    },
    rightBank: {
      missionaries: 1,
      cannibals: 1
    },
    boatIsOn: "right bank",
    goal: "all 3 missionaries and all 3 cannibals on the right bank",
    previousMove: "1 missionary + 1 cannibal from the left bank",
    statesVisitedSoFar: [
      {
        state: "3M 3C on the left, 0M 0C on the right, boat on the left",
        times: 1
      },
      {
        state: "2M 2C on the left, 1M 1C on the right, boat on the right",
        times: 1
      }
    ],
    legalMoves: {
      MOVE_1_MISSIONARY: {
        leadsTo: "3M 2C on the left, 0M 1C on the right, boat on the left",
        alreadyVisited: false,
        timesVisited: 0
      },
      MOVE_1_MISSIONARY_1_CANNIBAL: {
        leadsTo: "3M 3C on the left, 0M 0C on the right, boat on the left",
        alreadyVisited: true,
        timesVisited: 1
      }
    }
  },
  question: "Three missionaries and three cannibals must all cross a river from the left bank to the right bank. The boat carries one or two people and cannot cross empty. On either bank, if any missionaries are present, cannibals must never outnumber them. Here is the current state and the legal moves available right now. Which move should be made next? You are also told which states have already been visited.",
  actions: ["MOVE_1_MISSIONARY", "MOVE_1_MISSIONARY_1_CANNIBAL"],
  descriptions: {
    MOVE_1_MISSIONARY: "Take 1 missionary from the right bank to the left bank. Afterwards: left bank 3 missionaries and 2 cannibals; right bank 0 missionaries and 1 cannibals; boat on the left bank. You have not been in that state before.",
    MOVE_1_MISSIONARY_1_CANNIBAL: "Take 1 missionary + 1 cannibal from the right bank to the left bank. Afterwards: left bank 3 missionaries and 3 cannibals; right bank 0 missionaries and 0 cannibals; boat on the left bank. You have already been in that state 1 time."
  }
}
05 · Tic-Tac-Toe
Traditional

The whole game tree. Minimax searches every continuation to the end and knows the result of each move.

Jev

The board, its mark and the empty squares. At + tactics, one-move-ahead facts per square; at oracle, the perfect-play result per square.

Example payload, verbatim
{
  state: {
    youPlay: "X",
    opponentPlays: "O",
    board: ["O . .", ". X .", ". . ."],
    boardLegend: "rows top to bottom, '.' is empty",
    squares: {
      TOP_LEFT: "O",
      TOP_CENTER: "empty",
      TOP_RIGHT: "empty",
      MIDDLE_LEFT: "empty",
      CENTER: "X",
      MIDDLE_RIGHT: "empty",
      BOTTOM_LEFT: "empty",
      BOTTOM_CENTER: "empty",
      BOTTOM_RIGHT: "empty"
    },
    movesPlayedSoFar: 2
  },
  question: "You are playing tic-tac-toe. Players take turns placing their mark on an empty square of a 3×3 board; the first to get three in a row horizontally, vertically or diagonally wins, and a full board with no line is a draw. It is your turn. Which empty square should you take?",
  actions: [
    "TOP_CENTER",
    "TOP_RIGHT",
    "MIDDLE_LEFT",
    "MIDDLE_RIGHT",
    "BOTTOM_LEFT",
    "BOTTOM_CENTER",
    "BOTTOM_RIGHT"
  ],
  descriptions: {
    TOP_CENTER: "Place X on the top center square.",
    TOP_RIGHT: "Place X on the top right square.",
    MIDDLE_LEFT: "Place X on the middle left square.",
    MIDDLE_RIGHT: "Place X on the middle right square.",
    BOTTOM_LEFT: "Place X on the bottom left square.",
    BOTTOM_CENTER: "Place X on the bottom center square.",
    BOTTOM_RIGHT: "Place X on the bottom right square."
  }
}
Example payload · + tactics level
{
  state: {
    youPlay: "O",
    opponentPlays: "X",
    board: ["O . X", ". X .", ". . ."],
    boardLegend: "rows top to bottom, '.' is empty",
    squares: {
      TOP_LEFT: "O",
      TOP_CENTER: "empty",
      TOP_RIGHT: "X",
      MIDDLE_LEFT: "empty",
      CENTER: "X",
      MIDDLE_RIGHT: "empty",
      BOTTOM_LEFT: "empty",
      BOTTOM_CENTER: "empty",
      BOTTOM_RIGHT: "empty"
    },
    movesPlayedSoFar: 3
  },
  question: "You are playing tic-tac-toe. Players take turns placing their mark on an empty square of a 3×3 board; the first to get three in a row horizontally, vertically or diagonally wins, and a full board with no line is a draw. It is your turn. Which empty square should you take? Each option lists facts about what it does one move ahead.",
  actions: [
    "TOP_CENTER",
    "MIDDLE_LEFT",
    "MIDDLE_RIGHT",
    "BOTTOM_LEFT",
    "BOTTOM_CENTER",
    "BOTTOM_RIGHT"
  ],
  descriptions: {
    TOP_CENTER: "Place O on the top center square. After this, X can win on their next move.",
    MIDDLE_LEFT: "Place O on the middle left square. After this, X can win on their next move.",
    MIDDLE_RIGHT: "Place O on the middle right square. After this, X can win on their next move.",
    BOTTOM_LEFT: "Place O on the bottom left square. This blocks X from completing three in a row.",
    BOTTOM_CENTER: "Place O on the bottom center square. After this, X can win on their next move.",
    BOTTOM_RIGHT: "Place O on the bottom right square. After this, X can win on their next move."
  }
}
Exhibit 04 · 10 episodes per level

How much context does planning need?

Each level is still one local decision per move; only the information grows. Reference rows use no model at all: a random mover, with and without the same memory. Jev + search is the chess-engine split: code searches, Jev evaluates.
Solved within 30 stepsOptimal (11 moves)
Jev · State only
2.1 states visited
0/10
0/10
Jev · + Memory
12.3 states visited
10/10
7/10
Jev · + Lookahead
12 states visited
10/10
10/10
Jev · Oracle
12 states visited
10/10
10/10
Luna off · State only
4.5 states visited
0/10
0/10
Luna off · + Memory
12.8 states visited
10/10
2/10
Luna off · + Lookahead
12 states visited
10/10
10/10
Luna off · Oracle
12 states visited
10/10
10/10
Luna on · State only
11 states visited
8/10
3/10
Luna on · + Memory
12.2 states visited
10/10
9/10
Luna on · + Lookahead
12 states visited
10/10
10/10
Luna on · Oracle
12 states visited
10/10
10/10
Random mover
no model · 20,000 runs
6%
0%
Random + memory
avoids visited states
98%
33%
Traditional
15
States BFS expands
Finds the 11-move optimum. No model calls.
Jev
14
States Jev-guided search expands
Finds an 11-move plan with 14 Jev calls to score states.
Traditional
15
States a one-line heuristic expands
Same search, scoring states by how many people are across. No model.

Reading it. With the current state alone Jev locks into a two-state loop, no better than a random mover (6% solved). A record of visited states turns that into a solve, and most of the credit belongs to the memory: random moves that avoid visited states solve it 98% of the time. With lookahead facts Jev is optimal, but a no-model rule given the same facts is optimal 100% of the time, so at that level the facts do the planning.

The engine view. Search supplies the guarantee; Jev supplies the scores. Guided by Jev's scores the search expanded 14 states, against 15 for BFS and 15 for a one-line "people across" heuristic. On a 16-state puzzle that shows the split works, not that Jev adds efficiency. The 8-Puzzle exhibit tests the same idea on 181,440 states.

Exhibit 05 · 20 games per setting

Against a perfect opponent

Perfect play from the empty board is always a draw, so a loss is always someone's mistake. A blunder is a Jev move that turns a drawn (or won) position into a worse one under perfect play.
Not lost vs perfect playBlunder rateWon vs random
Jev · Board only
4/20
27%
12/20
Jev · + Tactics
14/20
9%
17/20
Jev · Oracle
20/20
0%
19/20
Luna off · Board only
7/20
19%
13/20
Luna off · + Tactics
16/20
4%
18/20
Luna off · Oracle
20/20
0%
16/20
Luna on · Board only
18/20
3%
16/20
Luna on · + Tactics
19/20
1%
19/20
Luna on · Oracle
20/20
0%
20/20
Minimax
always
0%
Rule + lookahead facts
no model · same facts as Jev
100%
100%

Reading it. The same shape as Missionaries. From the board alone Jev plays plausible tic-tac-toe, enough to beat a random player, but a perfect opponent punishes its oversights. One-move tactics computed by code change its blunder rate from 27% to 9%; with the perfect-play result per square it is 0%. Jev chooses well when the options are described well; it is not a search algorithm.

Exhibit 06 · 4 puzzles per depth

Jev as a search heuristic

A* on the 8-puzzle, recreating the textbook comparison of heuristics (AIMA 3rd ed., Fig. 3.29) with Jev added. Fewer nodes generated means a better guide. A* is only guaranteed to find the shortest solution if the heuristic never overestimates.
DepthNo heuristicMisplaced tilesManhattanJev, board onlyJev + featuresLuna, board onlyLuna + features
8
623
b* 2.06
39
b* 1.35
27
b* 1.26
33
b* 1.30
31
b* 1.30
50
b* 1.40
40
b* 1.35
12
4,109
b* 1.88
157
b* 1.37
43
b* 1.19
133
b* 1.33
69
b* 1.25
152
b* 1.33
1 not optimal
119
b* 1.32
16
28,813
b* 1.80
1,151
b* 1.44
226
b* 1.25
670
b* 1.39
1 over budget
408
b* 1.32
426
b* 1.32
1 over budget
224
b* 1.26
1 over budget

Mean nodes generated over 4 puzzles per depth; b* is AIMA's effective branching factor. Model runs stop after scoring 600 boards. Luna ran with reasoning switched off on this problem: at its default setting a single 8-puzzle call reasoned for more than three minutes. Its other problems used the default setting.

Jev estimates · board only
00551010151520202525true 16 · estimate ≈9 · 4 boardstrue 17 · estimate ≈8.5 · 4 boardstrue 15 · estimate ≈8.5 · 5 boardstrue 17 · estimate ≈12.5 · 13 boardstrue 17 · estimate ≈10.5 · 4 boardstrue 16 · estimate ≈11.5 · 8 boardstrue 14 · estimate ≈9.5 · 5 boardstrue 18 · estimate ≈12 · 9 boardstrue 18 · estimate ≈11 · 8 boardstrue 16 · estimate ≈12 · 10 boardstrue 18 · estimate ≈13.5 · 12 boardstrue 13 · estimate ≈7 · 2 boardstrue 14 · estimate ≈10.5 · 12 boardstrue 12 · estimate ≈7.5 · 2 boardstrue 11 · estimate ≈8.5 · 2 boardstrue 13 · estimate ≈9.5 · 4 boardstrue 14 · estimate ≈12 · 2 boardstrue 19 · estimate ≈7.5 · 1 boardtrue 20 · estimate ≈9.5 · 1 boardtrue 18 · estimate ≈8.5 · 3 boardstrue 19 · estimate ≈7 · 2 boardstrue 19 · estimate ≈11 · 4 boardstrue 18 · estimate ≈10 · 4 boardstrue 18 · estimate ≈14 · 8 boardstrue 12 · estimate ≈11 · 3 boardstrue 10 · estimate ≈8.5 · 2 boardstrue 19 · estimate ≈8 · 1 boardtrue 18 · estimate ≈10.5 · 3 boardstrue 18 · estimate ≈9.5 · 2 boardstrue 15 · estimate ≈11 · 12 boardstrue 14 · estimate ≈8 · 3 boardstrue 16 · estimate ≈13 · 9 boardstrue 13 · estimate ≈6.5 · 3 boardstrue 15 · estimate ≈9.5 · 7 boardstrue 14 · estimate ≈7 · 1 boardtrue 13 · estimate ≈8 · 2 boardstrue 17 · estimate ≈12 · 14 boardstrue 16 · estimate ≈10 · 5 boardstrue 16 · estimate ≈10.5 · 6 boardstrue 14 · estimate ≈12.5 · 7 boardstrue 14 · estimate ≈11 · 9 boardstrue 15 · estimate ≈12 · 7 boardstrue 9 · estimate ≈7.5 · 2 boardstrue 8 · estimate ≈9 · 1 boardtrue 10 · estimate ≈9 · 2 boardstrue 12 · estimate ≈10.5 · 6 boardstrue 16 · estimate ≈14.5 · 13 boardstrue 19 · estimate ≈13.5 · 8 boardstrue 16 · estimate ≈7.5 · 2 boardstrue 17 · estimate ≈10 · 5 boardstrue 16 · estimate ≈12.5 · 13 boardstrue 17 · estimate ≈13 · 13 boardstrue 16 · estimate ≈11 · 6 boardstrue 15 · estimate ≈10 · 3 boardstrue 17 · estimate ≈9.5 · 5 boardstrue 16 · estimate ≈15 · 7 boardstrue 16 · estimate ≈13.5 · 11 boardstrue 19 · estimate ≈11.5 · 7 boardstrue 20 · estimate ≈13.5 · 5 boardstrue 20 · estimate ≈15.5 · 1 boardtrue 17 · estimate ≈11.5 · 8 boardstrue 15 · estimate ≈11.5 · 8 boardstrue 15 · estimate ≈7.5 · 4 boardstrue 20 · estimate ≈11.5 · 4 boardstrue 20 · estimate ≈12.5 · 9 boardstrue 20 · estimate ≈15 · 4 boardstrue 17 · estimate ≈11 · 6 boardstrue 19 · estimate ≈12 · 7 boardstrue 15 · estimate ≈13.5 · 8 boardstrue 13 · estimate ≈8.5 · 2 boardstrue 14 · estimate ≈10 · 7 boardstrue 12 · estimate ≈9.5 · 4 boardstrue 15 · estimate ≈9 · 2 boardstrue 16 · estimate ≈9.5 · 4 boardstrue 20 · estimate ≈12 · 6 boardstrue 19 · estimate ≈12.5 · 4 boardstrue 21 · estimate ≈12 · 2 boardstrue 21 · estimate ≈10.5 · 4 boardstrue 22 · estimate ≈13.5 · 2 boardstrue 14 · estimate ≈13 · 7 boardstrue 13 · estimate ≈10 · 4 boardstrue 20 · estimate ≈14 · 10 boardstrue 22 · estimate ≈14.5 · 4 boardstrue 11 · estimate ≈11 · 5 boardstrue 7 · estimate ≈6.5 · 1 boardtrue 6 · estimate ≈11.5 · 1 boardtrue 8 · estimate ≈11.5 · 2 boardstrue 16 · estimate ≈14 · 4 boardstrue 15 · estimate ≈12.5 · 8 boardstrue 13 · estimate ≈11 · 5 boardstrue 18 · estimate ≈13 · 15 boardstrue 15 · estimate ≈10.5 · 3 boardstrue 12 · estimate ≈12.5 · 5 boardstrue 11 · estimate ≈9.5 · 3 boardstrue 11 · estimate ≈10 · 2 boardstrue 18 · estimate ≈16.5 · 1 boardtrue 20 · estimate ≈14.5 · 5 boardstrue 18 · estimate ≈15 · 9 boardstrue 21 · estimate ≈14 · 4 boardstrue 19 · estimate ≈14 · 4 boardstrue 13 · estimate ≈12.5 · 5 boardstrue 13 · estimate ≈12 · 5 boardstrue 10 · estimate ≈11.5 · 4 boardstrue 17 · estimate ≈14.5 · 4 boardstrue 18 · estimate ≈14.5 · 6 boardstrue 12 · estimate ≈11.5 · 2 boardstrue 14 · estimate ≈15.5 · 4 boardstrue 15 · estimate ≈14.5 · 4 boardstrue 19 · estimate ≈14.5 · 2 boardstrue 19 · estimate ≈13 · 9 boardstrue 10 · estimate ≈14.5 · 4 boardstrue 14 · estimate ≈14.5 · 4 boardstrue 21 · estimate ≈9.5 · 1 boardtrue 22 · estimate ≈12 · 2 boardstrue 16 · estimate ≈17 · 2 boardstrue 13 · estimate ≈11.5 · 2 boardstrue 12 · estimate ≈13 · 7 boardstrue 10 · estimate ≈13 · 1 boardtrue 18 · estimate ≈12.5 · 8 boardstrue 11 · estimate ≈12 · 3 boardstrue 15 · estimate ≈13 · 4 boardstrue 16 · estimate ≈16 · 5 boardstrue 20 · estimate ≈16 · 3 boardstrue 22 · estimate ≈15.5 · 2 boardstrue 14 · estimate ≈15 · 3 boardstrue 17 · estimate ≈15 · 2 boardstrue 22 · estimate ≈12.5 · 4 boardstrue 16 · estimate ≈8.5 · 1 boardtrue 18 · estimate ≈11.5 · 7 boardstrue 21 · estimate ≈13.5 · 6 boardstrue 22 · estimate ≈17 · 1 boardtrue 15 · estimate ≈15 · 7 boardstrue 15 · estimate ≈14 · 7 boardstrue 13 · estimate ≈14 · 6 boardstrue 20 · estimate ≈17 · 1 boardtrue 14 · estimate ≈14 · 7 boardstrue 20 · estimate ≈13 · 3 boardstrue 21 · estimate ≈13 · 4 boardstrue 17 · estimate ≈13.5 · 9 boardstrue 19 · estimate ≈15 · 6 boardstrue 22 · estimate ≈14 · 2 boardstrue 15 · estimate ≈7 · 1 boardtrue 21 · estimate ≈11 · 1 boardtrue 21 · estimate ≈14.5 · 1 boardtrue 23 · estimate ≈15 · 1 boardtrue 22 · estimate ≈16 · 2 boardstrue 15 · estimate ≈8 · 1 boardtrue 12 · estimate ≈13.5 · 3 boardstrue 14 · estimate ≈13.5 · 5 boardstrue 18 · estimate ≈15.5 · 3 boardstrue 21 · estimate ≈12.5 · 4 boardstrue 21 · estimate ≈11.5 · 2 boardstrue 22 · estimate ≈13 · 1 boardtrue 20 · estimate ≈11 · 4 boardstrue 13 · estimate ≈10.5 · 2 boardstrue 12 · estimate ≈15 · 4 boardstrue 11 · estimate ≈12.5 · 3 boardstrue 14 · estimate ≈11.5 · 2 boardstrue 22 · estimate ≈16.5 · 3 boardstrue 23 · estimate ≈13.5 · 2 boardstrue 23 · estimate ≈12.5 · 1 boardtrue 22 · estimate ≈9.5 · 1 boardtrue 23 · estimate ≈11.5 · 1 boardtrue 21 · estimate ≈8 · 1 boardtrue 17 · estimate ≈14 · 4 boardstrue 13 · estimate ≈14.5 · 3 boardstrue 19 · estimate ≈15.5 · 2 boardstrue 11 · estimate ≈10.5 · 2 boardstrue 9 · estimate ≈9 · 2 boardstrue 8 · estimate ≈10 · 3 boardstrue 12 · estimate ≈14 · 3 boardstrue 19 · estimate ≈10.5 · 4 boardstrue 16 · estimate ≈15.5 · 3 boardstrue 23 · estimate ≈15.5 · 1 boardtrue 10 · estimate ≈13.5 · 3 boardstrue 21 · estimate ≈15 · 1 boardtrue 13 · estimate ≈13 · 3 boardstrue 9 · estimate ≈13 · 4 boardstrue 7 · estimate ≈14 · 1 boardtrue 7 · estimate ≈13 · 3 boardstrue 5 · estimate ≈11 · 1 boardtrue 22 · estimate ≈15 · 1 boardtrue 14 · estimate ≈17.5 · 2 boardstrue 11 · estimate ≈13 · 2 boardstrue 17 · estimate ≈9 · 1 boardtrue 19 · estimate ≈9.5 · 1 boardtrue 21 · estimate ≈15.5 · 1 boardtrue 8 · estimate ≈10.5 · 2 boardstrue 7 · estimate ≈8.5 · 1 boardtrue 6 · estimate ≈11 · 3 boardstrue 10 · estimate ≈9.5 · 4 boardstrue 7 · estimate ≈12.5 · 6 boardstrue 5 · estimate ≈11.5 · 1 boardtrue 11 · estimate ≈7 · 1 boardtrue 13 · estimate ≈7.5 · 3 boardstrue 8 · estimate ≈13 · 1 boardtrue 4 · estimate ≈9.5 · 1 boardtrue 6 · estimate ≈15 · 1 boardtrue 3 · estimate ≈8 · 5 boardstrue 2 · estimate ≈6 · 5 boardstrue 4 · estimate ≈10.5 · 2 boardstrue 1 · estimate ≈3.5 · 8 boardstrue 3 · estimate ≈9 · 2 boardstrue 3 · estimate ≈6 · 6 boardstrue 2 · estimate ≈5 · 5 boardstrue 0 · estimate ≈3 · 7 boardstrue 8 · estimate ≈13.5 · 3 boardstrue 9 · estimate ≈14.5 · 6 boardstrue 8 · estimate ≈12 · 3 boardstrue 6 · estimate ≈10.5 · 1 boardtrue 7 · estimate ≈12 · 4 boardstrue 5 · estimate ≈7 · 3 boardstrue 7 · estimate ≈10.5 · 1 boardtrue 4 · estimate ≈8 · 4 boardstrue 6 · estimate ≈9.5 · 5 boardstrue 4 · estimate ≈11 · 5 boardstrue 3 · estimate ≈7.5 · 5 boardstrue 9 · estimate ≈13.5 · 6 boardstrue 8 · estimate ≈14 · 2 boardstrue 5 · estimate ≈8 · 5 boardstrue 7 · estimate ≈11 · 3 boardstrue 2 · estimate ≈6.5 · 3 boardstrue 2 · estimate ≈5.5 · 3 boardstrue 6 · estimate ≈12.5 · 2 boardstrue 9 · estimate ≈10 · 2 boardstrue 10 · estimate ≈10.5 · 1 boardtrue 8 · estimate ≈14.5 · 4 boardstrue 10 · estimate ≈14 · 3 boardstrue 11 · estimate ≈11.5 · 3 boardstrue 5 · estimate ≈9.5 · 2 boardstrue 4 · estimate ≈7.5 · 2 boardstrue 6 · estimate ≈8.5 · 2 boardstrue 6 · estimate ≈9 · 2 boardstrue 12 · estimate ≈17 · 1 boardtrue 13 · estimate ≈16 · 2 boardstrue 11 · estimate ≈15 · 2 boardstrue 12 · estimate ≈15.5 · 1 boardtrue 10 · estimate ≈15.5 · 1 boardtrue 14 · estimate ≈16 · 3 boardstrue 10 · estimate ≈15 · 2 boardstrue 15 · estimate ≈17 · 1 boardtrue 11 · estimate ≈13.5 · 2 boardstrue 12 · estimate ≈16 · 2 boardstrue 11 · estimate ≈14.5 · 2 boardstrue 16 · estimate ≈17.5 · 1 boardtrue 6 · estimate ≈10 · 2 boardstrue 5 · estimate ≈7.5 · 2 boardstrue 4 · estimate ≈8.5 · 1 boardtrue 4 · estimate ≈11.5 · 1 boardtrue 10 · estimate ≈11 · 4 boardstrue 17 · estimate ≈3.5 · 1 boardtrue 18 · estimate ≈7.5 · 1 boardtrue 18 · estimate ≈5.5 · 1 boardtrue 17 · estimate ≈6 · 1 boardtrue 13 · estimate ≈13.5 · 5 boardstrue 12 · estimate ≈12 · 2 boardstrue 12 · estimate ≈10 · 2 boardstrue 8 · estimate ≈12.5 · 3 boardstrue 7 · estimate ≈10 · 2 boardstrue 4 · estimate ≈9 · 1 boardtrue 9 · estimate ≈11.5 · 1 boardtrue 10 · estimate ≈10 · 1 boardtrue 3 · estimate ≈8.5 · 1 boardtrue 12 · estimate ≈16.5 · 3 boardstrue 9 · estimate ≈10.5 · 1 boardtrue 9 · estimate ≈8 · 1 boardtrue 10 · estimate ≈8 · 1 boardtrue 11 · estimate ≈9 · 2 boardstrue 12 · estimate ≈8 · 2 boardstrue 13 · estimate ≈6 · 1 boardtrue 14 · estimate ≈8.5 · 2 boardstrue 15 · estimate ≈16 · 1 boardtrue 14 · estimate ≈9 · 1 boardtrue 19 · estimate ≈9 · 1 boardtrue 15 · estimate ≈15.5 · 2 boardstrue 16 · estimate ≈16.5 · 2 boardstrue 18 · estimate ≈9 · 1 boardtrue 3 · estimate ≈7 · 1 boardtrue 0 · estimate ≈3.5 · 1 boardtrue 12 · estimate ≈14.5 · 1 boardtrue 14 · estimate ≈16.5 · 1 boardtrue 9 · estimate ≈14 · 1 boardtrue 10 · estimate ≈16 · 1 boardtrue 6 · estimate ≈13 · 1 boardtrue 8 · estimate ≈11 · 1 boardtrue 9 · estimate ≈8.5 · 1 boardtrue 13 · estimate ≈15 · 2 boardstrue 18 · estimate ≈16 · 3 boardstrue 16 · estimate ≈18 · 1 boardtrue 19 · estimate ≈16 · 1 boardtrue 8 · estimate ≈15.5 · 1 boardtrue 17 · estimate ≈15.5 · 2 boardstrue 15 · estimate ≈16.5 · 1 boardtrue 17 · estimate ≈16 · 1 boardtrue 6 · estimate ≈12 · 1 boardtrue moves to goalestimate
1045 boards · mean error 4.2 moves · 26% overestimated (above the dashed line) · larger circles are several boards
Jev estimates · + misplaced and Manhattan counts
00551010151520202525true 16 · estimate ≈7.5 · 16 boardstrue 17 · estimate ≈6.5 · 7 boardstrue 15 · estimate ≈6 · 6 boardstrue 17 · estimate ≈9.5 · 12 boardstrue 16 · estimate ≈8 · 8 boardstrue 14 · estimate ≈5.5 · 8 boardstrue 18 · estimate ≈8 · 4 boardstrue 18 · estimate ≈7.5 · 9 boardstrue 16 · estimate ≈8.5 · 7 boardstrue 18 · estimate ≈7 · 1 boardtrue 13 · estimate ≈5.5 · 5 boardstrue 14 · estimate ≈7 · 8 boardstrue 12 · estimate ≈5.5 · 1 boardtrue 11 · estimate ≈5 · 2 boardstrue 13 · estimate ≈6.5 · 2 boardstrue 19 · estimate ≈6 · 2 boardstrue 20 · estimate ≈7 · 2 boardstrue 12 · estimate ≈6.5 · 5 boardstrue 10 · estimate ≈6 · 1 boardtrue 15 · estimate ≈9 · 11 boardstrue 14 · estimate ≈7.5 · 4 boardstrue 18 · estimate ≈10 · 9 boardstrue 19 · estimate ≈6.5 · 3 boardstrue 17 · estimate ≈9 · 20 boardstrue 15 · estimate ≈6.5 · 6 boardstrue 19 · estimate ≈7 · 2 boardstrue 19 · estimate ≈10 · 8 boardstrue 9 · estimate ≈5.5 · 2 boardstrue 14 · estimate ≈6.5 · 6 boardstrue 16 · estimate ≈10 · 19 boardstrue 15 · estimate ≈8.5 · 3 boardstrue 15 · estimate ≈8 · 3 boardstrue 12 · estimate ≈8 · 1 boardtrue 15 · estimate ≈10 · 8 boardstrue 18 · estimate ≈9 · 6 boardstrue 19 · estimate ≈10.5 · 6 boardstrue 13 · estimate ≈7 · 4 boardstrue 8 · estimate ≈5 · 1 boardtrue 10 · estimate ≈7 · 4 boardstrue 16 · estimate ≈11 · 5 boardstrue 16 · estimate ≈10.5 · 7 boardstrue 18 · estimate ≈9.5 · 9 boardstrue 20 · estimate ≈11.5 · 4 boardstrue 20 · estimate ≈9.5 · 3 boardstrue 7 · estimate ≈6 · 3 boardstrue 13 · estimate ≈8 · 4 boardstrue 21 · estimate ≈10.5 · 7 boardstrue 16 · estimate ≈7 · 4 boardstrue 16 · estimate ≈5.5 · 2 boardstrue 15 · estimate ≈7 · 3 boardstrue 17 · estimate ≈7.5 · 6 boardstrue 20 · estimate ≈11 · 4 boardstrue 20 · estimate ≈13 · 5 boardstrue 15 · estimate ≈7.5 · 4 boardstrue 15 · estimate ≈5.5 · 2 boardstrue 19 · estimate ≈11 · 3 boardstrue 19 · estimate ≈9.5 · 5 boardstrue 17 · estimate ≈10.5 · 9 boardstrue 12 · estimate ≈8.5 · 7 boardstrue 14 · estimate ≈8.5 · 4 boardstrue 17 · estimate ≈10 · 18 boardstrue 11 · estimate ≈9.5 · 1 boardtrue 18 · estimate ≈10.5 · 7 boardstrue 18 · estimate ≈11 · 12 boardstrue 17 · estimate ≈7 · 2 boardstrue 18 · estimate ≈6.5 · 2 boardstrue 20 · estimate ≈10 · 3 boardstrue 15 · estimate ≈9.5 · 4 boardstrue 13 · estimate ≈8.5 · 4 boardstrue 14 · estimate ≈9 · 5 boardstrue 16 · estimate ≈6 · 3 boardstrue 11 · estimate ≈7 · 3 boardstrue 6 · estimate ≈8 · 4 boardstrue 8 · estimate ≈7.5 · 2 boardstrue 14 · estimate ≈11 · 4 boardstrue 14 · estimate ≈12 · 5 boardstrue 11 · estimate ≈9 · 4 boardstrue 11 · estimate ≈8 · 3 boardstrue 15 · estimate ≈11.5 · 2 boardstrue 14 · estimate ≈10.5 · 3 boardstrue 18 · estimate ≈12 · 10 boardstrue 21 · estimate ≈12 · 2 boardstrue 21 · estimate ≈9.5 · 2 boardstrue 22 · estimate ≈10.5 · 2 boardstrue 20 · estimate ≈12.5 · 5 boardstrue 18 · estimate ≈13 · 4 boardstrue 20 · estimate ≈10.5 · 5 boardstrue 22 · estimate ≈12 · 2 boardstrue 22 · estimate ≈11 · 1 boardtrue 14 · estimate ≈8 · 3 boardstrue 12 · estimate ≈9.5 · 3 boardstrue 21 · estimate ≈10 · 4 boardstrue 19 · estimate ≈13 · 2 boardstrue 22 · estimate ≈12.5 · 3 boardstrue 16 · estimate ≈9.5 · 3 boardstrue 19 · estimate ≈8.5 · 5 boardstrue 16 · estimate ≈9 · 6 boardstrue 20 · estimate ≈12 · 8 boardstrue 14 · estimate ≈9.5 · 4 boardstrue 16 · estimate ≈12 · 5 boardstrue 14 · estimate ≈6 · 3 boardstrue 13 · estimate ≈10.5 · 6 boardstrue 10 · estimate ≈10.5 · 2 boardstrue 12 · estimate ≈9 · 2 boardstrue 16 · estimate ≈11.5 · 3 boardstrue 16 · estimate ≈12.5 · 2 boardstrue 14 · estimate ≈13 · 2 boardstrue 13 · estimate ≈10 · 3 boardstrue 13 · estimate ≈9.5 · 2 boardstrue 18 · estimate ≈8.5 · 5 boardstrue 13 · estimate ≈9 · 7 boardstrue 12 · estimate ≈10 · 5 boardstrue 10 · estimate ≈8.5 · 2 boardstrue 10 · estimate ≈7.5 · 4 boardstrue 18 · estimate ≈11.5 · 6 boardstrue 13 · estimate ≈12 · 1 boardtrue 22 · estimate ≈13 · 2 boardstrue 21 · estimate ≈11 · 4 boardstrue 15 · estimate ≈4.5 · 1 boardtrue 11 · estimate ≈8.5 · 3 boardstrue 9 · estimate ≈6 · 2 boardstrue 8 · estimate ≈5.5 · 2 boardstrue 19 · estimate ≈12.5 · 1 boardtrue 19 · estimate ≈12 · 3 boardstrue 16 · estimate ≈13 · 3 boardstrue 22 · estimate ≈9.5 · 1 boardtrue 19 · estimate ≈11.5 · 2 boardstrue 13 · estimate ≈5 · 3 boardstrue 12 · estimate ≈6 · 1 boardtrue 21 · estimate ≈11.5 · 1 boardtrue 23 · estimate ≈12.5 · 1 boardtrue 15 · estimate ≈10.5 · 3 boardstrue 14 · estimate ≈10 · 3 boardstrue 20 · estimate ≈9 · 1 boardtrue 13 · estimate ≈6 · 3 boardstrue 18 · estimate ≈12.5 · 1 boardtrue 17 · estimate ≈8 · 2 boardstrue 22 · estimate ≈10 · 1 boardtrue 21 · estimate ≈8.5 · 1 boardtrue 18 · estimate ≈6 · 1 boardtrue 20 · estimate ≈8.5 · 2 boardstrue 9 · estimate ≈9 · 2 boardstrue 9 · estimate ≈10 · 6 boardstrue 19 · estimate ≈9 · 3 boardstrue 12 · estimate ≈11 · 5 boardstrue 14 · estimate ≈11.5 · 3 boardstrue 21 · estimate ≈12.5 · 1 boardtrue 20 · estimate ≈13.5 · 1 boardtrue 11 · estimate ≈7.5 · 1 boardtrue 12 · estimate ≈7.5 · 1 boardtrue 15 · estimate ≈13 · 3 boardstrue 12 · estimate ≈10.5 · 4 boardstrue 10 · estimate ≈8 · 1 boardtrue 9 · estimate ≈10.5 · 5 boardstrue 17 · estimate ≈11.5 · 1 boardtrue 17 · estimate ≈11 · 5 boardstrue 7 · estimate ≈9.5 · 11 boardstrue 5 · estimate ≈8 · 1 boardtrue 15 · estimate ≈11 · 2 boardstrue 17 · estimate ≈12 · 2 boardstrue 17 · estimate ≈12.5 · 1 boardtrue 9 · estimate ≈11 · 4 boardstrue 22 · estimate ≈11.5 · 1 boardtrue 17 · estimate ≈8.5 · 1 boardtrue 9 · estimate ≈9.5 · 2 boardstrue 21 · estimate ≈9 · 1 boardtrue 20 · estimate ≈7.5 · 2 boardstrue 20 · estimate ≈6.5 · 2 boardstrue 4 · estimate ≈6.5 · 7 boardstrue 6 · estimate ≈8.5 · 3 boardstrue 3 · estimate ≈4.5 · 11 boardstrue 2 · estimate ≈3.5 · 11 boardstrue 4 · estimate ≈7 · 4 boardstrue 1 · estimate ≈2 · 9 boardstrue 3 · estimate ≈3.5 · 6 boardstrue 2 · estimate ≈3 · 6 boardstrue 0 · estimate ≈1.5 · 7 boardstrue 17 · estimate ≈13 · 2 boardstrue 13 · estimate ≈7.5 · 2 boardstrue 18 · estimate ≈14 · 1 boardtrue 11 · estimate ≈11 · 3 boardstrue 19 · estimate ≈14 · 1 boardtrue 10 · estimate ≈9 · 3 boardstrue 16 · estimate ≈13.5 · 2 boardstrue 19 · estimate ≈15 · 1 boardtrue 8 · estimate ≈8 · 3 boardstrue 10 · estimate ≈5.5 · 3 boardstrue 12 · estimate ≈5 · 1 boardtrue 7 · estimate ≈10.5 · 2 boardstrue 5 · estimate ≈7 · 3 boardstrue 10 · estimate ≈11 · 1 boardtrue 8 · estimate ≈10.5 · 8 boardstrue 8 · estimate ≈9 · 2 boardstrue 6 · estimate ≈7.5 · 4 boardstrue 5 · estimate ≈6 · 6 boardstrue 7 · estimate ≈8.5 · 1 boardtrue 4 · estimate ≈6 · 7 boardstrue 6 · estimate ≈5.5 · 5 boardstrue 7 · estimate ≈6.5 · 5 boardstrue 3 · estimate ≈5 · 3 boardstrue 9 · estimate ≈11.5 · 1 boardtrue 6 · estimate ≈7 · 1 boardtrue 7 · estimate ≈9 · 4 boardstrue 10 · estimate ≈12 · 1 boardtrue 10 · estimate ≈9.5 · 2 boardstrue 9 · estimate ≈6.5 · 1 boardtrue 5 · estimate ≈7.5 · 1 boardtrue 3 · estimate ≈4 · 3 boardstrue 12 · estimate ≈13.5 · 1 boardtrue 13 · estimate ≈14.5 · 1 boardtrue 11 · estimate ≈13 · 1 boardtrue 10 · estimate ≈12.5 · 3 boardstrue 12 · estimate ≈13 · 2 boardstrue 11 · estimate ≈13.5 · 1 boardtrue 8 · estimate ≈10 · 2 boardstrue 2 · estimate ≈4 · 1 boardtrue 11 · estimate ≈10 · 2 boardstrue 12 · estimate ≈12 · 2 boardstrue 11 · estimate ≈10.5 · 1 boardtrue 18 · estimate ≈5.5 · 1 boardtrue 10 · estimate ≈10 · 1 boardtrue 8 · estimate ≈7 · 2 boardstrue 15 · estimate ≈12.5 · 1 boardtrue 9 · estimate ≈8 · 2 boardstrue 9 · estimate ≈8.5 · 2 boardstrue 13 · estimate ≈12.5 · 2 boardstrue 12 · estimate ≈12.5 · 4 boardstrue 13 · estimate ≈13.5 · 1 boardtrue 11 · estimate ≈12.5 · 2 boardstrue 11 · estimate ≈6.5 · 1 boardtrue 8 · estimate ≈8.5 · 1 boardtrue 13 · estimate ≈11 · 1 boardtrue 0 · estimate ≈2 · 2 boardstrue 13 · estimate ≈14 · 2 boardstrue 10 · estimate ≈13 · 1 boardtrue 14 · estimate ≈12.5 · 1 boardtrue 13 · estimate ≈11.5 · 1 boardtrue 6 · estimate ≈9 · 1 boardtrue 5 · estimate ≈5.5 · 1 boardtrue 5 · estimate ≈6.5 · 1 boardtrue moves to goalestimate
872 boards · mean error 5.5 moves · 19% overestimated (above the dashed line) · larger circles are several boards
Luna (reasoning off) estimates · board only
00551010151520202525true 16 · estimate ≈5 · 40 boardstrue 17 · estimate ≈2 · 24 boardstrue 15 · estimate ≈5 · 42 boardstrue 17 · estimate ≈5 · 47 boardstrue 18 · estimate ≈2 · 26 boardstrue 18 · estimate ≈5 · 35 boardstrue 17 · estimate ≈8 · 44 boardstrue 16 · estimate ≈11 · 19 boardstrue 14 · estimate ≈0 · 9 boardstrue 13 · estimate ≈2 · 30 boardstrue 14 · estimate ≈5 · 40 boardstrue 12 · estimate ≈2 · 23 boardstrue 11 · estimate ≈2 · 11 boardstrue 13 · estimate ≈8 · 27 boardstrue 13 · estimate ≈5 · 38 boardstrue 18 · estimate ≈8 · 35 boardstrue 16 · estimate ≈2 · 57 boardstrue 15 · estimate ≈2 · 35 boardstrue 15 · estimate ≈8 · 34 boardstrue 10 · estimate ≈5 · 16 boardstrue 14 · estimate ≈8 · 20 boardstrue 19 · estimate ≈2 · 11 boardstrue 20 · estimate ≈2 · 17 boardstrue 19 · estimate ≈5 · 29 boardstrue 19 · estimate ≈8 · 44 boardstrue 17 · estimate ≈11 · 19 boardstrue 14 · estimate ≈2 · 47 boardstrue 12 · estimate ≈8 · 10 boardstrue 15 · estimate ≈11 · 16 boardstrue 18 · estimate ≈11 · 30 boardstrue 18 · estimate ≈17 · 3 boardstrue 18 · estimate ≈0 · 11 boardstrue 17 · estimate ≈14 · 11 boardstrue 19 · estimate ≈11 · 16 boardstrue 9 · estimate ≈2 · 10 boardstrue 8 · estimate ≈5 · 8 boardstrue 15 · estimate ≈14 · 12 boardstrue 20 · estimate ≈8 · 23 boardstrue 20 · estimate ≈5 · 28 boardstrue 21 · estimate ≈14 · 13 boardstrue 21 · estimate ≈8 · 28 boardstrue 12 · estimate ≈5 · 24 boardstrue 16 · estimate ≈8 · 23 boardstrue 18 · estimate ≈14 · 18 boardstrue 19 · estimate ≈14 · 21 boardstrue 16 · estimate ≈0 · 5 boardstrue 19 · estimate ≈17 · 6 boardstrue 7 · estimate ≈5 · 11 boardstrue 11 · estimate ≈11 · 8 boardstrue 11 · estimate ≈8 · 13 boardstrue 11 · estimate ≈5 · 28 boardstrue 17 · estimate ≈27 · 2 boardstrue 14 · estimate ≈27 · 7 boardstrue 16 · estimate ≈14 · 6 boardstrue 21 · estimate ≈2 · 1 boardtrue 22 · estimate ≈5 · 5 boardstrue 22 · estimate ≈2 · 3 boardstrue 21 · estimate ≈11 · 13 boardstrue 23 · estimate ≈8 · 8 boardstrue 14 · estimate ≈17 · 1 boardtrue 18 · estimate ≈27 · 8 boardstrue 12 · estimate ≈27 · 2 boardstrue 6 · estimate ≈5 · 15 boardstrue 8 · estimate ≈2 · 12 boardstrue 9 · estimate ≈8 · 13 boardstrue 13 · estimate ≈11 · 10 boardstrue 10 · estimate ≈2 · 14 boardstrue 9 · estimate ≈5 · 13 boardstrue 20 · estimate ≈11 · 21 boardstrue 7 · estimate ≈11 · 3 boardstrue 7 · estimate ≈8 · 9 boardstrue 5 · estimate ≈5 · 14 boardstrue 12 · estimate ≈11 · 12 boardstrue 21 · estimate ≈5 · 13 boardstrue 22 · estimate ≈11 · 14 boardstrue 22 · estimate ≈8 · 15 boardstrue 10 · estimate ≈0 · 7 boardstrue 9 · estimate ≈11 · 5 boardstrue 14 · estimate ≈11 · 16 boardstrue 20 · estimate ≈17 · 5 boardstrue 20 · estimate ≈14 · 11 boardstrue 22 · estimate ≈14 · 2 boardstrue 4 · estimate ≈2 · 21 boardstrue 3 · estimate ≈2 · 22 boardstrue 2 · estimate ≈2 · 18 boardstrue 4 · estimate ≈5 · 6 boardstrue 1 · estimate ≈2 · 9 boardstrue 16 · estimate ≈27 · 8 boardstrue 13 · estimate ≈14 · 6 boardstrue 19 · estimate ≈20 · 3 boardstrue 14 · estimate ≈23 · 2 boardstrue 0 · estimate ≈0 · 9 boardstrue 16 · estimate ≈17 · 3 boardstrue 21 · estimate ≈20 · 1 boardtrue 19 · estimate ≈27 · 3 boardstrue 8 · estimate ≈8 · 7 boardstrue 17 · estimate ≈17 · 3 boardstrue 7 · estimate ≈2 · 3 boardstrue 10 · estimate ≈20 · 1 boardtrue 10 · estimate ≈17 · 2 boardstrue 6 · estimate ≈2 · 8 boardstrue 5 · estimate ≈2 · 5 boardstrue 10 · estimate ≈11 · 10 boardstrue 10 · estimate ≈8 · 8 boardstrue 13 · estimate ≈17 · 1 boardtrue 9 · estimate ≈20 · 1 boardtrue 7 · estimate ≈14 · 1 boardtrue 6 · estimate ≈8 · 1 boardtrue 6 · estimate ≈0 · 3 boardstrue 8 · estimate ≈11 · 2 boardstrue 8 · estimate ≈14 · 1 boardtrue 10 · estimate ≈14 · 3 boardstrue 10 · estimate ≈27 · 1 boardtrue 5 · estimate ≈8 · 2 boardstrue 3 · estimate ≈5 · 2 boardstrue 12 · estimate ≈14 · 4 boardstrue 13 · estimate ≈27 · 1 boardtrue 11 · estimate ≈20 · 1 boardtrue 9 · estimate ≈14 · 1 boardtrue 12 · estimate ≈0 · 2 boardstrue 11 · estimate ≈14 · 3 boardstrue 16 · estimate ≈20 · 1 boardtrue 16 · estimate ≈23 · 2 boardstrue 8 · estimate ≈27 · 1 boardtrue 15 · estimate ≈27 · 2 boardstrue 15 · estimate ≈17 · 1 boardtrue 11 · estimate ≈27 · 1 boardtrue 14 · estimate ≈14 · 3 boardstrue 15 · estimate ≈23 · 2 boardstrue 13 · estimate ≈23 · 1 boardtrue 20 · estimate ≈0 · 2 boardstrue 22 · estimate ≈17 · 2 boardstrue 21 · estimate ≈27 · 2 boardstrue 20 · estimate ≈27 · 3 boardstrue 18 · estimate ≈23 · 1 boardtrue 23 · estimate ≈11 · 4 boardstrue 22 · estimate ≈0 · 3 boardstrue 24 · estimate ≈11 · 3 boardstrue 22 · estimate ≈27 · 2 boardstrue 20 · estimate ≈20 · 1 boardtrue 23 · estimate ≈14 · 4 boardstrue 20 · estimate ≈23 · 1 boardtrue 24 · estimate ≈14 · 4 boardstrue 24 · estimate ≈17 · 1 boardtrue 24 · estimate ≈23 · 2 boardstrue 23 · estimate ≈5 · 1 boardtrue 24 · estimate ≈8 · 1 boardtrue 22 · estimate ≈20 · 1 boardtrue 23 · estimate ≈27 · 1 boardtrue 22 · estimate ≈23 · 1 boardtrue moves to goalestimate
1717 boards · mean error 8.5 moves · 8% overestimated (above the dashed line) · larger circles are several boards
Luna (reasoning off) estimates · + misplaced and Manhattan counts
00551010151520202525true 16 · estimate ≈5 · 52 boardstrue 17 · estimate ≈5 · 46 boardstrue 15 · estimate ≈5 · 40 boardstrue 17 · estimate ≈8 · 40 boardstrue 18 · estimate ≈5 · 50 boardstrue 14 · estimate ≈2 · 13 boardstrue 13 · estimate ≈5 · 20 boardstrue 18 · estimate ≈2 · 3 boardstrue 19 · estimate ≈5 · 19 boardstrue 14 · estimate ≈5 · 30 boardstrue 12 · estimate ≈5 · 20 boardstrue 16 · estimate ≈8 · 31 boardstrue 20 · estimate ≈5 · 14 boardstrue 11 · estimate ≈2 · 4 boardstrue 12 · estimate ≈2 · 4 boardstrue 10 · estimate ≈2 · 10 boardstrue 9 · estimate ≈5 · 18 boardstrue 19 · estimate ≈8 · 39 boardstrue 19 · estimate ≈2 · 3 boardstrue 18 · estimate ≈8 · 26 boardstrue 16 · estimate ≈2 · 7 boardstrue 20 · estimate ≈8 · 23 boardstrue 14 · estimate ≈8 · 16 boardstrue 13 · estimate ≈2 · 3 boardstrue 15 · estimate ≈2 · 2 boardstrue 15 · estimate ≈8 · 36 boardstrue 11 · estimate ≈5 · 15 boardstrue 21 · estimate ≈11 · 3 boardstrue 8 · estimate ≈5 · 26 boardstrue 10 · estimate ≈5 · 23 boardstrue 18 · estimate ≈11 · 14 boardstrue 10 · estimate ≈8 · 17 boardstrue 19 · estimate ≈14 · 2 boardstrue 21 · estimate ≈14 · 2 boardstrue 21 · estimate ≈8 · 23 boardstrue 7 · estimate ≈5 · 29 boardstrue 11 · estimate ≈8 · 24 boardstrue 13 · estimate ≈8 · 27 boardstrue 20 · estimate ≈11 · 14 boardstrue 18 · estimate ≈14 · 2 boardstrue 16 · estimate ≈11 · 16 boardstrue 22 · estimate ≈11 · 4 boardstrue 22 · estimate ≈8 · 7 boardstrue 12 · estimate ≈8 · 20 boardstrue 6 · estimate ≈5 · 23 boardstrue 8 · estimate ≈2 · 5 boardstrue 9 · estimate ≈8 · 22 boardstrue 5 · estimate ≈5 · 20 boardstrue 20 · estimate ≈14 · 3 boardstrue 22 · estimate ≈5 · 2 boardstrue 21 · estimate ≈5 · 4 boardstrue 14 · estimate ≈11 · 3 boardstrue 8 · estimate ≈8 · 13 boardstrue 20 · estimate ≈2 · 2 boardstrue 4 · estimate ≈2 · 15 boardstrue 3 · estimate ≈2 · 24 boardstrue 2 · estimate ≈2 · 19 boardstrue 4 · estimate ≈5 · 17 boardstrue 1 · estimate ≈2 · 10 boardstrue 17 · estimate ≈11 · 4 boardstrue 19 · estimate ≈11 · 3 boardstrue 23 · estimate ≈8 · 1 boardtrue 17 · estimate ≈14 · 3 boardstrue 15 · estimate ≈14 · 2 boardstrue 9 · estimate ≈2 · 2 boardstrue 7 · estimate ≈8 · 9 boardstrue 0 · estimate ≈0 · 9 boardstrue 6 · estimate ≈2 · 13 boardstrue 10 · estimate ≈14 · 1 boardtrue 10 · estimate ≈11 · 4 boardstrue 9 · estimate ≈11 · 1 boardtrue 5 · estimate ≈2 · 7 boardstrue 12 · estimate ≈11 · 6 boardstrue 13 · estimate ≈14 · 6 boardstrue 11 · estimate ≈11 · 6 boardstrue 7 · estimate ≈2 · 2 boardstrue 20 · estimate ≈0 · 2 boardstrue 13 · estimate ≈11 · 5 boardstrue 15 · estimate ≈17 · 2 boardstrue 15 · estimate ≈11 · 1 boardtrue 22 · estimate ≈14 · 2 boardstrue moves to goalestimate
1110 boards · mean error 7.1 moves · 5% overestimated (above the dashed line) · larger circles are several boards

Reading it. The diagonal is a perfect heuristic; the shaded area below it is where an estimate is safe for A*. Given the Manhattan distance as a feature, the question is whether Jev adds anything beyond it: compare the Manhattan column with the Jev + features column. In “Jev decides” mode (no search), it solved 0 of 10 puzzles with the board alone and 1 of 10 with a memory of visited boards (30-move limit, 10 moves optimal).

Exhibit 07 · in scope

Tagging real text

Part-of-speech tagging on 200 sentences of the Universal Dependencies English Web Treebank, scored against human annotation. Every word is a 49-way Choice: the kind of many-label, one-second judgment TypeSafe says Jev is built for.
Jev
2593 tokens
92.9%
Luna, reasoning off
2593 tokens
93.6%
Luna, reasoning on
2593 tokens
94.5%
HMM + Viterbi
textbook tagger, trained on EWT
92.3%
Most frequent tag
per-word lookup
88.4%

Jev's most common mistakes (gold → Jev): NN→NNP ×17 · ,→: ×11 · VBN→JJ ×9 · NN→JJ ×9 · IN→TO ×8 · NNP→NN ×7. Same sentences, same tokenization, for every tagger.

Exhibit 06 · in scope

Ranking boards, code searches

TypeSafe's recommended pattern: code runs the search and the model only picks which next board looks closest to solved. The search is iterative widening: a beam of width 1, then 2, 3, 5, 8, 13, reusing rankings already made, up to 300 distinct rankings per puzzle. A perfect ranker and an uninformative one bracket the result, so the score measures the ranker, not the search.
RankerSolvedDepth 8Depth 12Depth 16OptimalMean rankings
Perfect ranker
true distance · ceiling
30 / 3089100%10/1010/1010/1030/309
Manhattan ranker
no model
29 / 308399%10/1010/109/1018/3064
Jev16 / 303670%8/105/103/109/30183
Luna, reasoning off12 / 302558%7/105/100/108/30173
All options equal
no information · floor
1 / 30117%1/100/100/101/30181

10 seeded puzzles per optimal depth. The perfect, Manhattan and uninformative rows use no model and are identical across runs. Luna returns a pick without probabilities, so its ranking gives 90% to its pick. Earlier versions of this page used a single width-3 beam, which cannot back up after one poor ranking; widening was adopted because it lets a good ranker succeed (the perfect ranker solves every puzzle).

Exhibit 02 · 5 presets + 40 random scenarios

Where Jev and the tree disagree

For each preset: the tree's rule, Jev's P(wait), and three controls. Tree-only removes kids and ratings; if Jev then agrees with the tree, the disagreement was about information. Repeat and reversed measure noise and option-order bias.
ScenarioTreeJevP(wait)Luna offLuna onTree-onlyRepeatReversedStrongest factor readings
Date night
WAIT
WaitEstimate 30–60, Alternate yes, Fri/Sat yes → wait
LEAVE
disagrees
36%WAIT
tree-only: leave
WAIT13% leave42%45%
Weather argues for staying no 8%
A good alternative exists yes 87%
Family with kids
WAIT
WaitEstimate 30–60, Alternate yes, Fri/Sat yes → wait
LEAVE
disagrees
10%LEAVELEAVE1% leave10%7%
Can wait comfortably no 5%
Weather argues for staying no 7%
Tourist in rain
WAIT
WaitEstimate 10–30, Hungry yes, Alternate no → wait
WAIT68%WAITWAIT16% leave70%64%
Can wait comfortably no 7%
Hard on someone in the party no 19%
Very hungry
LEAVE
WaitEstimate 30–60, Alternate yes, Fri/Sat no → leave
LEAVE9%LEAVELEAVE0% leave11%12%
Can wait comfortably no 6%
Weather argues for staying no 8%
Great restaurant / long wait
LEAVE
WaitEstimate >60 → leave
LEAVE36%WAIT
tree-only: leave
LEAVE5% leave37%37%
Weather argues for staying no 8%
A good alternative exists yes 86%

Reading it. Removing kids and ratings lowers Jev's P(wait) in 5 of 5 scenarios. In none of the disagreements does the reduced state move Jev onto the tree's side: Jev weighs the tree's own attributes differently, not just more of them. Open any scenario on the exhibit for the full rationale and single-change counterfactuals.

40 seeded random scenarios
  • Jev: agrees with the tree on 30 of 40 (6086%); waits in 22, the tree in 24.
  • Luna off: agrees with the tree on 31 of 40 (6288%); waits in 29, the tree in 24.
  • Luna on: agrees with the tree on 33 of 40 (6891%); waits in 29, the tree in 24.

Every input drawn independently, so some combinations are unusual. Agreement with a 1995 tree, not accuracy.

Exhibit 03 · 50 worlds

Same beliefs, different agents

Every agent reads the same exact risk estimates. The hand-written agent also plans paths across the known map; the models see their neighbours only.
AgentEscaped with goldNo gold, aliveDiedMean score
Traditional29 / 504471%183512
Jev23 / 503360%1710245
Luna off16 / 502146%331293
Luna on23 / 503360%1611228
Every world (50)
WorldTraditionalScoreJevScoreLuna offScoreLuna onScore
AIMA Fig. 7.2 escaped with gold988 escaped with gold979 escaped with gold981 escaped with gold981
Seed 1000 retreated-1 escaped with gold985 escaped with gold985 escaped with gold985
Seed 1001 retreated-1 retreated-30 retreated-1 escaped with gold992
Seed 1002 escaped with gold984 retreated-20 retreated-3 fell into a pit-1017
Seed 1003 escaped with gold994 retreated-11 retreated-11 retreated-11
Seed 1004 retreated-1 escaped with gold990 retreated-1 escaped with gold994
Seed 1005 escaped with gold994 escaped with gold996 escaped with gold996 escaped with gold996
Seed 1006 escaped with gold994 retreated-20 retreated-7 escaped with gold967
Seed 1007 retreated-1 fell into a pit-1001 retreated-1 fell into a pit-1001
Seed 1008 escaped with gold986 escaped with gold990 escaped with gold992 escaped with gold992
Seed 1009 escaped with gold988 escaped with gold994 retreated-28 escaped with gold963
Seed 1010 escaped with gold992 retreated-22 fell into a pit-1017 escaped with gold965
Seed 1011 escaped with gold978 escaped with gold969 retreated-3 retreated-3
Seed 1012 escaped with gold994 escaped with gold992 retreated-13 retreated-3
Seed 1013 retreated-1 retreated-41 retreated-1 retreated-5
Seed 1014 retreated-1 retreated-41 retreated-1 retreated-36
Seed 1015 escaped with gold990 escaped with gold992 escaped with gold988 fell into a pit-1006
Seed 1016 escaped with gold990 escaped with gold996 escaped with gold994 escaped with gold996
Seed 1017 retreated-1 escaped with gold985 escaped with gold985 escaped with gold985
Seed 1018 escaped with gold982 retreated-28 escaped with gold988 escaped with gold990
Seed 1019 retreated-1 retreated-39 retreated-1 fell into a pit-1015
Seed 1020 retreated-1 fell into a pit-1001 retreated-1 fell into a pit-1001
Seed 1021 escaped with gold980 retreated-28 escaped with gold990 escaped with gold990
Seed 1022 escaped with gold990 escaped with gold996 escaped with gold996 escaped with gold996
Seed 1023 escaped with gold992 retreated-28 escaped with gold971 escaped with gold961
Seed 1024 escaped with gold990 escaped with gold992 escaped with gold994 escaped with gold994
Seed 1025 fell into a pit-1004 fell into a pit-1006 retreated-3 retreated-3
Seed 1026 fell into a pit-1004 retreated-5 retreated-7 retreated-3
Seed 1027 escaped with gold988 retreated-20 retreated-3 fell into a pit-1002
Seed 1028 retreated-1 retreated-41 retreated-1 escaped with gold994
Seed 1029 retreated-1 escaped with gold996 retreated-1 escaped with gold996
Seed 1030 retreated-1 fell into a pit-1001 retreated-1 fell into a pit-1001
Seed 1031 escaped with gold996 escaped with gold994 retreated-3 retreated-3
Seed 1032 escaped with gold980 escaped with gold977 escaped with gold984 escaped with gold973
Seed 1033 fell into a pit-1010 retreated-13 retreated-3 retreated-3
Seed 1034 escaped with gold990 escaped with gold996 escaped with gold996 escaped with gold996
Seed 1035 retreated-1 fell into a pit-1001 retreated-1 fell into a pit-1001
Seed 1036 escaped with gold992 fell into a pit-1008 escaped with gold975 retreated-28
Seed 1037 escaped with gold990 escaped with gold994 retreated-20 retreated-34
Seed 1038 retreated-1 escaped with gold979 retreated-1 retreated-26
Seed 1039 escaped with gold984 escaped with gold986 retreated-20 escaped with gold975
Seed 1040 retreated-1 escaped with gold994 retreated-1 escaped with gold994
Seed 1041 escaped with gold980 retreated-5 retreated-3 retreated-3
Seed 1042 retreated-1 fell into a pit-1001 retreated-1 fell into a pit-1001
Seed 1043 escaped with gold982 escaped with gold971 escaped with gold988 escaped with gold984
Seed 1044 escaped with gold988 fell into a pit-1018 retreated-5 retreated-18
Seed 1045 escaped with gold971 retreated-30 retreated-3 retreated-11
Seed 1046 retreated-1 fell into a pit-1001 retreated-1 fell into a pit-1001
Seed 1047 escaped with gold994 escaped with gold992 retreated-3 retreated-3
Seed 1048 retreated-1 fell into a pit-1001 retreated-1 fell into a pit-1001
Robustness

Consistency checks

Two properties a decision component should have, measured on the five restaurant scenarios.
Jev
2 pts
Mean change when the identical call is repeated
Jev is not perfectly deterministic. Differences this size are noise.
Jev
4 pts
Mean change when the options are listed in reverse
Option-order bias. Small compared with the effect of real inputs.
Context

Among other Jev demos

How this project compares with the public Jev demos and evaluations posted on X in the first three days after launch.
PracticePublic demos doing itThis project
Compares Jev against something
Any baseline: another model, embeddings, a heuristic or an algorithm.
5 / 16
yes
a textbook algorithm on every problem
Compares against a non-LLM method
Embeddings, a rule, a heuristic or a search algorithm.
1 / 16
yes
rules, AIMA tree, minimax, BFS, random movers
Varies what Jev is told
Runs the same task with more or less context, to separate model from inputs.
0 / 16
yes
context ladders on Missionaries and Tic-Tac-Toe
Code does the search, Jev chooses
Lookahead or filtering is done by ordinary code and handed to Jev.
2 / 16
yes
engine mode, tactics and lookahead levels
Reports where Jev fails
Losses, errors or limits, not only successes.
5 / 16
yes
0/3 state-only, losses to minimax
Quantified results
Counts, rates or scores rather than a single anecdote.
7 / 16
yes
small samples: 3–8 per setting
Code or live demo to reproduce
A repository or a playable page.
5 / 16
yes
live exhibits + npm run benchmark
Compares against an LLM
The same task given to a general language model.
2 / 16
yes
GPT-5.6 Luna on every problem, same payloads
Reports cost per decision
Dollars per call or per task.
3 / 16
yes
tokens and list-price USD for every call
Real-world task
Production-like data rather than a teaching problem.
7 / 16
no
deliberately textbook problems

Where this stands. Most public Jev posts are single demonstrations of speed and cost on a real task. They are good at showing Jev is fast and cheap, and most do not compare it with anything. This project is the only one found that tests Jev against textbook algorithms and varies what Jev is told. Only Mario (+ lookahead) matches our finding that code-supplied lookahead is what makes the model look capable, and its author plans the obvious next test: Jev versus a cheap heuristic on the same simulations.

Where others are ahead. They use real data (moderation, payments, pull requests, retrieval). This project deliberately uses textbook problems, and its samples are small.

DemoKindWhat it does (as the post describes it)BaselineFailures shown
This projectGame / control · Classification · PlanningFive textbook AI problems, each against a classical solution, with context ladders and recorded runs.classical algorithmsyes
@TheINAOG
Sep 17
Game / controlSuper Mario Bros. 1-1 cleared for $0.04. RAM parsed into structured state; code simulates controller macros and filters predicted deaths; Jev picks among the survivors.heuristic (planned)yes
@kanemama_
Sep 17
Game / controlChess: Jev picks moves probabilistically from the legal move set. Playable.none stated
@DineshDataAI
Sep 17
Game / controlTetris: Jev returns a typed placement (rotation, column, drop); the game engine validates it.none statedyes
@akafukusou
Sep 17
RetrievalEvidence retrieval on 34 QASPER questions against pgvector + OpenAI embeddings.embeddingsyes
@frankiedigiac
Sep 17
ClassificationReal-time payment-transaction screening compared with Claude Haiku 4.5 on the same streams.LLM
@iamMrDuncan
Sep 18
EvaluationOpen benchmark of a small local model (Needle 3) against Jev, Qwen 27B and OSS 120B.local models
@redp314
Sep 17
Dev toolingPull-request review: one call returns 14 typed checks (secrets, SQL injection, …) as probabilities.LLM
@tamarajtran
Sep 18
Dev toolingInstant context compaction: score every tool call and drop the irrelevant ones, instead of summarising.none stated
@altryne
Sep 18
Dev toolingThe compaction idea as a Claude plugin: a session of nearly 1M tokens cut to 86K in about a second.none stated
@kacpersinilo
Sep 17
ClassificationComment moderation: 20 comments in 1.48 s of API time for $0.00044.none stated
@stevekrouse
Sep 16
PlaygroundLive playground for trying Jev questions.none stated
@waynesutton
Sep 17
Playgroundaskjev.ai: ask anything; instead of answering, Jev judges.none stated
@sydneyrunkle
Sep 18
Dev toolingArticle on using Jev as the evaluation step inside an agent harness loop.none stated
@justALEXWORTEGA
Sep 17
CritiqueClaims an MLP head on a 4B open model behaves like Jev; weights published.local models
@HiromTeachesAI
Sep 16
CritiqueThread arguing Jev is a fast classifier, not a chat-model competitor, and questioning the launch claims.none statedyes
@Salman_Bareesh
Sep 17
CritiqueTried wiring Jev into a scoring pipeline; reports SDK errors and a billing wall before any results.none statedyes

Collected from X search on 18 September 2026, three days after Jev launched: 16 posts from the top results for Jev queries (demos, evaluations, critiques). Not exhaustive; launch announcements, questions and reposts are left out. Each row reflects only what the post itself says, so a practice counted as missing may exist elsewhere.

Method

How this was run

  • Every Jev decision is a single Choice question over the legal actions only. Probabilities are as returned.
  • Questions were written once. Missionaries context levels add information computed by ordinary code; none adds instructions about strategy.
  • Traditional solutions: the AIMA reflex rule, the AIMA restaurant tree, a hand-written probabilistic Wumpus agent, and breadth-first search.
  • The LLM column: scripts/benchmark.ts sends GPT-5.6 Luna exactly the payload Jev receives, through the same server-side task builders, with the answer constrained to the legal labels by a JSON schema. Luna returns a choice but no probabilities; none are invented. Results are included above.
  • Samples: 10 vacuum dirt schedules, 5 restaurant presets plus 40 random scenarios, 50 Wumpus worlds, 10 Missionaries episodes per level, 20 tic-tac-toe games per setting, 30 8-puzzles for ranking, 200 POS sentences (2593 words). Rates carry 95% Wilson intervals; POS accuracy a bootstrap interval over sentences. Jev is nearly deterministic, so repeated episodes of the same start add little information: Missionaries and Tic-Tac-Toe show consistency more than breadth.