Files
claude b351bc2d55
ci / verify (push) Successful in 48s
Stage 4: upgrades, and a quartermaster to spend them at
Every level banks one choice. Choices queue, and are spent at an NPC in the
hub: walk to it, press E, take one of three weighted options. Seven upgrades,
all data — split shot, glass cannon, spread, sniper, doubleshot, poison,
eraser — and SimWorld gained no per-upgrade branch to run any of them.

The four ambiguities in the brief were settled with the user first, since
each changes what gets written:

  damage      base x (1 + sum additive) x product multiplicative. The flat
              +5% every upgrade carries, spread's -10%, doubleshot's -50%
              and glass cannon's +100% pool; sniper multiplies on top, so
              two snipers is 4x and not +200%.
  glass       half the LEVELLED maximum, multiplying if taken twice, so the
              price does not fade to a rounding error by level 15.
  poison      independent stacks, not a refresh.
  split       +/-45 degrees from the original heading.

Independent poison stacks sound expensive and are not: every dose lasts the
same number of ticks, so doses expire in the order they were added, the
pending expiries are a plain FIFO, and PoisonTrack only ever looks at its
front. O(1) per actor per tick however many are live.

Stats are derived from the upgrade list and never stored, the way level is
derived from experience -- a saved stat cannot disagree with the upgrades
that produced it. Upgrade riders (split charges, poison, erase chance) travel
on the bullet instead, because a shot in flight has to keep what it was fired
with rather than gaining Poison because the shooter just took it.

Two invariants this collided with, both now pinned:

  - bullet speed gained a ceiling. Wall collision samples once per tick, so
    anything over a tile per tick tunnels; two snipers asked for 2480 u/s
    against a 1920 threshold, and a tunnelling bullet looks like a bullet.
  - BULLET_INTEREST_RADIUS rose to 2900, because an upgraded player shot is
    now the longest-travelling bullet in the game. test_interest measured
    the worst case from static content, which upgrades quietly invalidated.

Choosing is intent checked three ways: a choice must be owed, the index must
name one of the three options the SERVER put on the table, and the player
must be standing at the NPC. The offer is rolled once and persisted, so
closing the screen is not a reroll and neither is a crash.

tools/diag_upgrades.tscn covers level -> banked choice -> refused in a
dungeon and refused across the room -> taken at the NPC -> new stats ->
on disk. Bots never walk to the quartermaster, so the smoke test cannot.

Known gap recorded in the roadmap: at PLAYER_BULLET_DAMAGE = 6, the +5% the
first upgrade carries rounds back to 6 and visibly does nothing. It comes out
right in aggregate, but the fix is a balance edit across content.gd and so is
the user's call.

check.sh clean, 357 tests, SMOKE PASS (18 assertions), all four diagnostics
green.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-06 15:34:58 +02:00

448 lines
23 KiB
Markdown
Raw Permalink Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
# Netcode
## Model
Authoritative dedicated server, clients send input only, with client-side
prediction and reconciliation for the local player.
```
client server (60 Hz)
| |
|-- InputFrame (60 Hz, x3) ---->| queue_input(): validate, drop, sort
| move, aim, 5 button bits, | SimWorld.step(): move, AI, emitters,
| inventory slot | integrate bullets, resolve hits, items
|<-- events (reliable, ch 3) ---| bullet spawns/despawns, hits, deaths,
| | item pickups/uses/drops
|<-- snapshot (unrel., ch 2) ---| players, enemies, boss, ground loot,
| | your own inventory @ 20 Hz
| |
predict locally, reconcile authoritative for everything
```
The client has no message that expresses a position, a hit, damage taken, a
completed escape, or an item it now owns. This is deliberate and stronger than
validating such messages after the fact: there is no code path to exploit, only
intent to interpret.
Item actions are worth noting as the newest thing to resist becoming a message
of its own. Using or dropping an item is a request that reaches the server as
two button bits and a slot number on the ordinary input frame, and the server
decides what — if anything — happened. See *Item actions ride the input frame*
below.
## Why bullets are not replicated as state
A bullet hell has hundreds of live bullets — this project peaks at ~350 with one
boss and four players. Sending a position per bullet per snapshot would be
roughly 350 × 8 bytes × 20 Hz ≈ 56 KB/s per client and would grow with the
pattern density, which is the one thing you want freedom to increase.
Instead the server sends **one spawn event per bullet, once**: uid, position,
velocity, radius, lifetime, kind, team, acceleration, turn rate (~36 bytes).
Both sides then run the identical integration in `BulletPool.step()`. A bullet
that expires or leaves the arena dies on both sides with no traffic at all.
The only bullets that need a despawn message are the ones killed early by a hit,
because that is the one outcome a client cannot derive.
Cost at the same peak: a few KB/s, dominated by whatever the patterns are
actually firing rather than by how much is in the air.
`tests/integration/test_replica_parity.gd` is the test that keeps this honest.
It runs the server and a replica side by side for 900 ticks of a boss fight —
once in-process, once with every event encoded and decoded through the real wire
format — and asserts that every live bullet matches within 0.01 px (1 px through
float32 on the wire). If that test fails, the bullets on screen are no longer
the bullets that can kill you.
Note what this does **not** require: lockstep determinism. Clients simulate
bullets for display only. The server never reads a client's bullet state, so
float drift between machines is a cosmetic concern, not a correctness one.
## Hit validation
All in `SimWorld._resolve_bullet_hits()` and `_resolve_contact_damage()`, server
only, circle-vs-circle:
- **Enemy bullet → player:** skipped entirely while the player has i-frames, so
a dense pattern costs one hit, not thirty. The bullet is consumed and a
`BULLET_DESPAWN` goes out.
- **Player bullet → boss, then enemies:** first overlap consumes the bullet.
Boss damage is scaled by the current phase's `damage_taken_mult`.
- **Enemy body → player:** contact damage, same i-frame gate.
There is no lag compensation and no rewind. For a bullet hell that is the right
call: what the player is dodging is their own position against a bullet field,
and rewinding the world to a shooter's view would mean a player who dodged on
their screen still gets hit. The cost is that a high-ping player's shots at a
moving enemy land where the enemy *was* on the server. Enemies are large and
slow relative to bullet speed, so this reads as latency rather than unfairness.
If it becomes a complaint, the fix is to lag-compensate **player bullets against
enemies only**, never enemy bullets against players.
## Input timing, and a bug worth remembering
The client numbers its input frames `INPUT_TARGET_LEAD` ticks ahead of the
server's tick; the server accepts anything within `INPUT_MAX_LEAD` and re-syncs
nothing itself. **These two numbers must be chosen together.**
They originally were not — the server dropped inputs past a lead of 12, while
the client only corrected itself past 16. A client whose lead drifted into 1316
therefore had *every* input silently rejected, while believing its own timing
was fine. The server, receiving nothing, coasted on the last input it had
(`held_input`) and then stopped; the client kept predicting forward. The ship
and the authoritative position separated permanently, the reconciler fought it
every snapshot, and the player got shoved around. It needed two independent
clocks to drift apart, so it appeared only after minutes of play — and never
recovered, because nothing in the loop could notice.
Two defences now:
1. `INPUT_MAX_LEAD` (40) is far wider than the client's correction band
(`INPUT_LEAD_MIN` 3 … `INPUT_LEAD_MAX` 20), so the client always re-syncs
long before the server starts refusing anything. `test_input_lead.gd`
asserts that ordering directly, so narrowing the window fails a test rather
than shipping.
2. An **ack-stall detector**: if `last_input_tick` does not advance across
`INPUT_ACK_STALL_LIMIT` snapshots, the client concludes its numbering is
outside the window and hard re-syncs. Lead alone cannot detect this, because
a wrong lead looks perfectly normal from the client's side. This is the
backstop that makes the whole class of failure self-healing regardless of
cause.
`tools/diag_prediction.gd` injects a +14 tick drift mid-run and asserts the gap
returns to normal; it exits non-zero if it does not.
### Every exit runs the same channel
There is deliberately no "I am leaving cleanly" message. The escape channel is
keyed on the socket closing, so all four ways out converge on it:
| Route | Path |
| --- | --- |
| Hold F | `BTN_ESCAPE` -> channel |
| Menu -> "Return to hub" | synthesises a held `BTN_ESCAPE` -> channel |
| Menu -> "Disconnect to menu" | `Net.shutdown()` -> socket close -> `linkdead` -> channel |
| Killing the process / pulling the cable | socket close -> `linkdead` -> channel |
The last two are the same server-side code path, and `tools/smoke.sh` asserts
both: one bot is SIGKILLed mid-dungeon, another calls the same `Net.shutdown()`
the menu button calls. Adding a clean-leave message that skipped the channel
would break that test, which is the point of having it.
The menu says so out loud when you are in a dungeon -- the mechanic only reads
as fair if the player knows the cost before clicking.
## Maps are streamed, never sent
A dungeon's geometry reaches a client as **chunks around that client's own
player**, and nothing else. In particular the server does *not* send the
generation seed, even though maps are a pure function of (kind, seed, depth) and
sending two integers would be far cheaper: a client holding the seed can
regenerate the entire dungeon locally, which is a complete map hack requiring no
skill. Streaming bounds what a modified client can know to roughly where its
player has physically been.
`MAP_STREAM_RADIUS` (900u) is deliberately wider than `FOG_VIEW_RADIUS` (460u).
The client needs real geometry it cannot see, because it predicts its own
movement against walls and simulates bullets that die on them. The accepted
trade is that a cheater sees somewhat further than the fog shows -- what they
cannot get is the floor plan.
Two consequences fall out of partial knowledge:
- **Wall deaths are announced.** A bullet dying to its lifetime or by leaving
the map is derivable from the spawn event and the map's dimensions, so it
costs nothing. A bullet dying against a *wall* is not derivable by a client
that has not been streamed that wall, so the server sends an explicit
`BULLET_DESPAWN`. Getting the order of those two checks wrong is easy and was
caught by a test: out-of-bounds tiles read as `WALL` by design, so testing
geometry before bounds reports every bullet that merely left the map as a
wall kill.
- **Unknown tiles are treated as empty**, not solid, so an unstreamed region can
never wrongly stop a prediction. The stream radius keeps far enough ahead of
the player that this never decides anything visible.
Hard fog is therefore a *rendering* rule, not a secrecy mechanism. The secrecy
lives in what the server declines to send: terrain outside the stream radius,
and (next stage) actors outside it.
## Actors are scoped per peer
Snapshots are encoded **once per peer**, not once and broadcast: an actor
further than `ACTOR_INTEREST_RADIUS` (800u) from a player is never sent to that
player. Before this, fog was hiding enemies the client had already been handed,
which is no defence at all against a modified client.
Measured cost of encoding four filtered snapshots instead of one shared:
95.6µs against 24.9µs, which amortised over the 3-tick snapshot interval is
0.032 ms/tick — 0.2% of the frame budget.
Two rules fall out of it:
- **The observer's own player record is never filtered**, however far outside
the radius the arithmetic puts it. The client reconciles its prediction
against that record; dropping it would break the player's own movement rather
than merely hide someone.
- **Bullet despawns are not filtered, only spawns.** A client that was told
about a bullet must always be told it died, or it keeps a phantom until the
lifetime expires.
Bullet spawns use a much wider radius (`BULLET_INTEREST_RADIUS`, 2200u) than
actors, because the failure modes are not symmetric. An enemy appearing at the
edge of sight is cosmetic; a bullet withheld at spawn that later flies into view
is invisible damage. The floor is *longest bullet travel + fog radius* — 1500 +
460 for the Warden's "Collapse" snipe — and `test_interest.gd` recomputes that
from the live content, so adding a faster or longer-lived bullet fails a test
rather than producing bullets that wink into existence.
## Loot has two visibilities, and one of them is an interest rule
Ground loot rides the snapshot rather than an event stream: items do not move,
so re-sending them 20 times a second costs almost nothing and a lost packet
costs nothing at all — which a spawn-once event could not claim.
Each item is either **world-shared** (`SimLoot.owner_peer == 0`) or
**player-instanced** (owned by exactly one peer). The instanced kind is filtered
in `NetCodec.encode_snapshot`, next to the actor interest radius and for the
same reason: a peer is never *told* that another player's copy exists, so a
modified client has nothing to reveal. Hiding it in the UI would have been the
same class of mistake as relying on fog to hide enemies.
Two consequences:
- **An instanced item leaves with its owner.** `SimWorld.remove_player` deletes
loot owned by the departing peer. Nobody else can see or take it, so leaving
it behind would be an invisible entity the instance carries until it closes.
- **Anything dropped becomes world-shared**, whatever it was before. That is
what makes dropping worth having.
The Warden's Ration exists to keep this path honest: it is dropped instanced on
every boss kill, so the filter runs in every real fight rather than only in
tests.
## Which dungeon you enter is a position, not a request
The hub holds a list of portals, each bound to a dungeon id. Pressing interact
resolves the player's **server-side** position through `SimWorld.portal_at()`,
and the resulting `PORTAL_USED` event carries which dungeon that entrance opens.
There is no client message that names a dungeon, and there must not be: one
would let any client ask for the Proving Grounds' loot rate — ten times the drop
chance — and bring the results back to the hub. The same reasoning as everywhere
else here; the difference between "the server checks your request" and "there is
no request" is the whole model.
`enter_instance` carries the portal list (position plus a dungeon index, via
`NetCodec.encode_portals`) and the id of the dungeon you have arrived in. The
client needs the latter to scale the boss's health ceiling the same way the
server did, or the boss bar would sit near empty for an entire easy fight.
## Item actions ride the input frame
`InputFrame` carries `BTN_USE`, `BTN_DROP` and a slot byte (10 bytes total, up
from 9). The alternative — a reliable `c_use_item(slot)` RPC — would have needed
its own replay guard, its own rate limit, and its own ordering story against the
movement on the same tick. The input stream already has all three.
The one thing it does not give for free is edge detection. The client repeats
its last few frames every tick (that redundancy is what covers a dropped packet)
and a starved server coasts on the last frame it holds, so a level-triggered
read would spend four items in four ticks. `SimPlayer.prev_buttons` and
`prev_slot` hold the edge; the slot is part of it, so tapping 2 while 1 is still
held is a second action rather than a swallowed one. Movement and fire stay
level-triggered — holding them is exactly what you mean.
Pickup shares `BTN_INTERACT` with the dungeon portal. Loot wins when both are in
reach, but only on a tick where something was actually taken, so a full bag
cannot leave a player standing on the portal unable to use it.
## Upgrades change what a shot is, and none of it is client-side
A player's damage, fire rate, bullet speed, shot count and every rider
(splitting, poison, erasure) come from `PlayerStats`, rebuilt on the server from
the upgrade ids the character holds. The client builds its own copy for the
choice screen's summary, and that copy decides nothing — it exists so the panel
can show real numbers instead of promises.
Three things fall out of it worth writing down:
- **The riders live on the bullet, not on the shooter.** A shot already in
flight keeps what it was fired with, so taking Poison mid-flight cannot reach
back and poison it. They stay zero on a replica, which never resolves a hit
and has no use for them.
- **Choosing an upgrade is intent, checked three ways.** The client sends an
index; the server checks a choice is owed, that the index names one of the
three options *it* put on the table, and that the player is standing at the
hub NPC. The index selects from a list the client did not write.
- **The offer is server state.** Rolled once, stored on the character, and
persisted — so closing the screen is not a reroll, and neither is a crash.
## No contact damage
Nothing hurts you by touching it. Every threat is a bullet you can see and
dodge, which is the contract the genre runs on; an enemy that damages you for
occupying the same space is an unavoidable, unreadable hit.
The Stalker used to be exactly that -- it walked at you and dealt contact
damage. It now carries a point-blank shotgun instead: five pellets, 62 degrees,
and an 18-tick lifetime that gives it about 78px of reach, so it still has to
close the distance and still leaves nothing lingering in the arena.
`tests/unit/test_content.gd` asserts that every hostile enemy has at least one
emitter, so a new enemy cannot quietly reintroduce the mechanic.
## No invulnerability frames
A hit grants no immunity — every bullet that touches you deals its damage. In a
bullet hell the wall *is* the threat, and i-frames invert that: you are punished
for the first bullet of a pattern and gifted the next thirty, which makes dense
patterns *safer* than sparse ones. Measured cost: a stationary player survives
~13.6s of the Warden's opening phase, ~17.5s while drifting aimlessly.
Arrival protection (`SPAWN_GRACE_TICKS`) is the sole exception, and it is a
transition, not a combat mechanic.
## Input validation
`SimWorld.queue_input()` is the single audit point. It drops:
| Case | Guard |
| --- | --- |
| Replay of a consumed tick | `f.tick <= p.last_input_tick` |
| Stale input | `f.tick < tick - INPUT_MAX_AGE` |
| Input claiming the future | `f.tick > tick + INPUT_MAX_LEAD` |
| Flood | queue capped at `INPUT_MAX_AGE`, oldest dropped |
| Leaving the hub early after death | `RESPAWN_LOCKOUT_TICKS`, server-side |
| Item action repeated by a held key | edge-triggered against `prev_buttons` / `prev_slot` |
| Inventory slot index out of range | `SimPlayer.take_slot` answers "nothing" |
| Upgrade choice made anywhere but the hub NPC | `SimWorld.at_upgrade_npc`, server-side position |
| Upgrade index outside the server's own offer | `CharacterStore.take_upgrade` refuses |
The respawn lockout is worth calling out: the HUD disables its button for the
same three seconds, but that is presentation. A client that ignores its own UI
and holds the bit down still waits, because the check lives in `SimWorld`.
Beyond that, the wire format itself constrains the cheat surface: the move
vector is two signed bytes at 1/100 precision, and `Movement.sanitize_move()`
clamps to the unit disc, so no packet can express extra speed. Fire rate is a
server-side cooldown; holding the button every tick fires at exactly the same
rate as tapping it correctly.
A client that stops sending coasts on its last input for `INPUT_MAX_AGE` ticks
and then stops, so a dropped connection does not leave a player sliding.
## Leaving a run, and why disconnecting is not an escape
The emergency escape is a **one-second server-owned channel** that damage does
*not* interrupt. That combination is deliberate and the two halves depend on
each other.
An interruptible channel sounds like the right kind of risk until you follow it
through: a player about to die under fire can never finish the channel, so
killing the game process becomes strictly better than using the button. The
escape hatch turns into the exploit.
So a dropped connection runs the same channel. On disconnect the peer is removed
from the send list but its player **stays in the world** (`SimPlayer.linkdead`),
treated as holding the escape button down, still fully killable, for the same
second everyone else pays. Only when that resolves is the peer forgotten. Pulling
the plug is therefore never cheaper than pressing the key, and there is no
timing window where it is.
Two details that are easy to get wrong and are covered by tests:
- An instance must not close while a linkdead body is still resolving
(`Instance.has_linkdead()`), or a solo player's drop would delete their body
on the next tick and hand back the exact exploit.
- Dying still ends the channel. That is the one thing the escape cannot beat,
and it is what keeps the second a real risk rather than a formality.
`tools/smoke.sh` SIGKILLs a bot mid-dungeon and asserts the server channels it
out rather than dropping it instantly.
## Arriving in a run
Entering a dungeon grants `SimConfig.SPAWN_GRACE_TICKS` (2s) of **arrival
protection**: invulnerable *and* unable to shoot. Both halves matter —
invulnerability alone would make the spawn point a free firing position. It is
flagged in the snapshot (`Protocol.F_SPAWN_GRACE`) so every client can draw it
on every protected ship, not just its own.
## Prediction and reconciliation
`ClientRuntime` keeps every unacknowledged `InputFrame`. Each snapshot echoes
`last_input_tick` per player; on receipt the client discards acknowledged
frames, takes the server's authoritative position, and replays the rest through
the same `Movement.step_player()` the server used. Divergence over 24 px snaps;
smaller errors lerp at 0.3 so ordinary jitter does not read as rubber-banding.
`ClientRuntime` sets `process_physics_priority = -10` so it samples and sends
input *before* `ServerRuntime` ticks. This matters only on a listen server,
where both live in one process: with the default tree order the server ran
first, so input sampled on frame N was not consumed until frame N+1, leaving the
drawn ship a permanent one tick (4px at 240 u/s) ahead of the authoritative one
— and bullets, which spawn at the authoritative position, visibly trailed behind
the ship. `tools/diag_prediction.gd` measures this gap; it is currently ~0.1px
mean, down from a constant 4.00px.
The other half of that fix is `SimConfig.PLAYER_MUZZLE_OFFSET`, derived from the
*visual* radius rather than the hitbox. A remote client still draws its ship a
tick or two ahead of the server, and that margin is what keeps bullets emerging
from the nose rather than the middle of the sprite.
Remote players, enemies and the boss are interpolated between the last two
snapshots. Bullets are not interpolated — they are simulated, and a spawn event
carries the server tick it was generated at so the client fast-forwards the
bullet by the packet's flight time (`BulletPool.advance_slot()`) instead of
popping it in at the muzzle a round-trip late.
## Channels
ENet channels keep a burst of reliable bullet events from head-of-line blocking
the snapshot stream:
| Channel | Traffic | Mode |
| --- | --- | --- |
| 1 | handshake, instance transitions | reliable |
| 2 | snapshots | unreliable, newest wins |
| 3 | bullet spawns/despawns, hits | reliable ordered |
| 4 | client input | unreliable ordered |
`Protocol.VERSION` is checked at handshake and mismatched clients are rejected
with a reason, rather than desyncing later.
## Instances
One server process hosts the lobby hub plus every concurrent dungeon. An
`Instance` is a `SimWorld` and a peer list; messages are addressed to the peers
of one instance, so a player never receives traffic for a world they are not in.
At ~0.24 ms/tick per instance, the limit is bandwidth and player count long
before it is CPU.
Transfers (portal in, escape out, dungeon cleared) are driven by server-only
events that `NetCodec` refuses to serialise, so the decision cannot leak to a
client as something it might try to send back.
## The listen server
"Host and play" runs a `ServerRuntime` and a `ClientRuntime` in one process. The
client is peer 1 and goes through the identical code path — it still sends input
and still learns outcomes from snapshots. Only the transport is short-circuited,
in `Net.send_*` / `Net.send_input`, which check for the local peer and call the
handler directly instead of going through ENet. No simulation code knows the
difference, so a bug cannot hide in a "host-only" branch.
## Known gaps
- No lag compensation (see above — deliberate, with a stated escape hatch).
- No snapshot delta compression. Full state at 20 Hz is fine at this actor
count; delta-encoding against the last acknowledged snapshot is the next step
if enemy counts grow.
- No encryption or authentication. `ENetMultiplayerPeer` supports DTLS; wire it
up before exposing a server to the internet.
- Bullet spawn events are per-bullet. Sending *pattern* events instead (emitter
id + tick + seed, clients regenerate) would cut bullet traffic by roughly the
count of each emitter's shot. Worth doing only if bandwidth becomes the
binding constraint — it couples the client to emitter behaviour, which the
current design deliberately avoids.