| Commit message (Collapse) | Author |
|
- sort [-r] [-n]: everything in RAM, which bounds it - a task has 8 KiB and a
file caps at 3 KiB, so 96 lines of 2400 bytes with room for the stack.
Insertion sort over an index; at that size the simple thing is fast and
costs far less ROM. `uniq -c f | sort -n -r` finally works.
NB the key needs a scratch SLOT of its own: gt() compares table entries, so
reusing slot i as scratch overwrites the entries being shifted into it.
- cmp: reports the first differing byte, silent when the files match. There
was no way to check a copy on the machine - you had to haul both to a host.
- strings [-n N]: pairs with xxd now that /bin/<prog> is readable. It finds
"gbos sm83 (Game Boy Color)" inside /bin/uname.
- du [dir]: 256-byte blocks, the unit df and the filesystem use. Bytes
overflow - /bin alone is 52 x 16 KiB, which wrapped a 16-bit counter and
printed 32768. Recurses by re-opendir'ing, since the kernel keeps only one
"directory being listed" cursor.
- time CMD: 1/64 s resolution from the tick counter. It rebuilds the child's
command line at $A000 - which is its own argument area, so everything is
copied to the stack first.
|
|
The terminal understood SGR and nothing else, so there was no way to erase the
screen - printing 18 newlines only scrolls the history up out of sight.
term_erase_all blanks the buffer, resets the colours, homes the cursor and
repaints; the CSI parser routes 'J' to it, and clear(1) is four putc calls.
|
|
gbfs has carried I_NLINK since the beginning and nothing ever incremented it,
so a file could only have one name. SYS_LINK adds a second directory entry
pointing at the same inode: the data is stored once, and both names read it.
That makes rm's job different. It now decrements the link count and only frees
the inode and its blocks when the LAST name goes - removing one of two links
used to free blocks the other still pointed at.
Refused, with the reasons that matter here: linking a directory (it would make
a cycle the tree walkers cannot survive), an existing name, and anything under
/bin, which is ROM.
The syscall takes a request block because the trap needs HL for its dispatch
table - and ln copies both paths to the stack first, since that block is a
static living at $A000, on top of the command line it is reading.
|
|
/# ls -l /bin > b
/# cut -d \s -f 1 b | uniq -c
52 -
- touch FILE...: gbfs has no timestamps (an MBC5 cart has no RTC), so touch
does the half that means something here - make an empty file. O_APPEND is
exactly right: it creates, and it does not truncate what is already there.
- tr [-d] SET1 [SET2]: ranges (a-z), -d to delete, and a short SET2 padded
with its last character, like tr(1). ASCII only - a 128-entry map.
- rev, uniq [-c]: one line each, from a file or stdin like the other filters.
Two things the tests caught. A bare "-" is a SET, not a flag, so `tr . -`
has to skip flag-matching only when a letter follows the dash. And uniq
compared an empty "line" against the last run at EOF and flushed a phantom
blank one - a stray "1" in uniq -c output.
tr and cut also take \s \t \n \r escapes, because the shell has no quoting:
`-d ' '` arrives as three tokens, so a space or tab delimiter was simply
untypeable. Shell quoting would be the better fix; this makes the tools
usable today without touching the parser.
Banks 49-53; 10 program slots left in the 1 MB ROM.
|
|
The commands only existed in the kernel's program table, so the only way to
find out what the machine could do was to already know (or read the source).
They are now a directory in the tree:
/# ls / /# xxd -l 16 /bin/uname
bin 0000: cd 6b 43 06 00 0e 00 f7 .kC.....
readme 0008: 18 fe 47 0e 03 f7 c9 21 ..G....!
/# ls -l /bin | head -2
- 1 16384 worker
/bin is synthetic, not disk: dir_find resolves "bin" in the root to BIN_INO
and names inside it through the NameTable, so `ls /bin` lists exactly what the
shell can run - one table, one truth, no second list to keep in sync. The root
listing carries "bin" as a synthetic first entry so it is discoverable by
walking the tree.
A program image is inode BIN_FILE|progid: it stats as a 16 KiB file (a program
IS its ROM bank) and reads through rom_getb, which maps that bank, takes the
byte, and maps the CALLER'S text bank back before returning - the caller is
itself executing from $4000, so leaving the wrong bank mapped would return
into another program's code. Verified byte-for-byte against the host's view of
the blob. The tree is read-only: create/remove/write inside /bin all fail.
xxd grows -l (stop after N bytes), because a program image is 2048 dump lines
- more than a pipe's temp file can hold, let alone the screen.
Two traps worth remembering, both found the hard way here: `and BIN_FILE`
destroys A, so the inode must be reloaded before inode_ptr (missing it once
made ls -l stat inode 0 and call every file an empty non-directory); and ls
now copies its directory argument to the stack first, because stat()'s request
block is a static living at $A000 - on top of the command line it was reading.
|
|
Eight bytes per line rather than the usual sixteen: offset + 8 hex pairs +
the ASCII column is 38 of the terminal's 40 columns, where 16 would wrap
every line and make the dump unreadable. -c overrides it.
No -s (skip): there is no seek syscall, so it streams from the start - the
same reason tail(1) has to keep a ring.
/# xxd h
0000: 48 65 6c 6c 6f 20 67 62 Hello gb
0008: 6f 73 0d 0a os..
blk(1) could already peek at raw disk blocks; nothing could look inside a
file. Verified against known bytes, from stdin, with -c, through a pipe, and
on a missing file.
|
|
Filling in the core utilities the box was missing:
- ls -l: type, links and size, now that stat() exists. The screen is 40
columns, so it is "d 1 256 name", not the GNU sprawl. Sizes were
previously unknowable except one file at a time with wc -c.
- df: blocks and inodes, used and free, on the 32 KiB cart disk.
- mv: no rename syscall, so it is cp + rm - and the copy has to finish before
the source goes, or a failure loses the file.
- tail [-n N]: no seek syscall, so it streams the input through a ring of N
lines rather than jumping to the end. N is capped by that ring.
- tee: the missing half of a pipeline - `cat big | tee copy | wc -l`.
- sleep: the thing every script needs; the kernel counts 1/64 s ticks.
- help: lists every command the shell knows, walked out of the kernel's own
name table. The box previously had no way to tell you what it could do.
- poweroff: the only safe way to stop a gbos that has data on it. The cart's
battery RAM is written when the emulator exits cleanly, so killing it loses
the whole filesystem - it comes back "formatted". Found while testing /rc.
|
|
Four things userland had no way to ask for:
- SYS_STAT: type, link count, size and inode for a path. The trap needs HL
for its dispatch table, so a call taking both a path and an output buffer
passes a request block - the shape SYS_NET already established. usr/stat.h
wraps it, header-only like sock.h, because the block must live in the
caller's RAM.
- SYS_FSSTAT: blocks/inodes, total and free. The bitmaps are already cached
write-through in fs_bmbuf, so this is the popcount boot_fs already prints.
- SYS_ISATTY: is this process's stdin the console? The shell needs it to tell
an interactive session from a script.
- O_APPEND: open a file positioned at EOF. There is no seek syscall, so this
is the only way to add to a file - without it nothing running ON the Game
Boy can build a multi-line file, shell scripts included.
Together these back ls -l, df, and `>>`.
|
|
cursor's
Reported as "the nickname w00t renders with its first and last letter a
different colour". It was neither the nick hash nor the Fano palette tables:
the colour buffer was right, the tile pixels were right, and the tilemap
ATTRIBUTE - which selects the palette - was stale.
render_tile writes pixels and refreshes the attr cache, but only
update_cursor_attr ever pushed an attribute into the tilemap, for the tile
under the cursor. A batched write (sys_write, i.e. every puts()) renders its
whole dirty span through render_tile and then updates exactly one attribute,
so the rest of the line kept whatever palette its row was created with -
palette 0, whose colours are white, red and cyan. Hence: ordinary white text
was always fine, red and cyan were fine by luck, and green/yellow/blue/
magenta silently rendered as red or cyan. In "<w00t>" the brackets are
putc'd (attribute written, w and t correct) while the nick is puts'd, so the
middle "00" kept the stale palette and came out red.
update_tile_attr now does that work for any (row, tile col) and term_write_end
calls it for every tile it renders. update_cursor_attr is a thin wrapper.
NB the wrapper MUST stay immediately after render_cursor_tile: that path falls
through rather than calling, and render_tile clobbers B/C, so the row/column
have to be re-read from wCurRow/wCurCol. Putting update_tile_attr there first
had it inherit garbage and scribble attributes over unrelated rows - visible
as yellow text in the boot log.
Verified by decoding recorded frames (control-socket "record", which needs a
display run - --headless produces no frames): ansi(1)'s word line now renders
red/green/yellow/blue/magenta/cyan/white instead of red/red/cyan/red/cyan/
cyan/white, the per-character bar and RAINBOW! still cycle correctly, and
<w00t> is uniformly blue between white brackets (was blue/red/red/blue).
Pre-existing: the same capture reproduces on c9f61fc, before this session.
|
|
/# ircd &
host$ irssi -c 10.0.0.2 /join #gbos
other GB$ irc 10.0.0.2 gb2 /join #gbos
Implements the subset a real client needs to get in and talk: NICK, USER,
PING/PONG, JOIN, PART, PRIVMSG/NOTICE (to a channel or to a nick), NAMES and
QUIT, with the 001/004/375/372/376 numerics clients wait for on registration,
353/366 on join, and 421 for anything else. One channel per client keeps the
bookkeeping in fixed-size tables; this is a handheld with 8 KiB of task RAM.
MAXCL is 3, bounded by the kernel socket table (listener + client each take a
slot, the rest left for other programs). A fourth client is refused and the
daemon keeps serving - which is how the ~30s close() stall turned up.
Every event is logged to the LCD, so the Game Boy shows its own server
traffic: "+ alice", "alice joined", "<alice> hello".
Verified with real clients over the hub: two and three simultaneous hosts
registering, joining, channel relay (sender not echoed to itself), private
messages by nick, PING/PONG, unknown-command numerics, QUIT propagation, the
overflow refusal, and a second Game Boy running gbos's own irc(1) in the same
channel as a laptop client - messages crossing both ways between them.
|
|
A server could only ever hold ONE connection, which makes an ircd pointless:
listen() turned the listener into the connection, and net_find_tcp demuxed on
local port alone, so two clients on 6667 were indistinguishable.
- net_find_tcp now matches an established socket on the full four-tuple
(local port + peer IP + peer port) and only falls back to a LISTENing
socket when none matches.
- A SYN at a listener no longer consumes it: tcp_spawn_conn clones a new
socket (local port + owning pid, so exit/kill still reclaims it) and the
listener keeps listening. accept() returns that socket; closing it leaves
the listener alone. A full table drops the SYN, and the client's retransmit
is taken once a slot frees.
- tcp_peer_is_current preserves BC: net_find_tcp scans with the socket index
there, and the compare needs BC for the rx buffer base. Without this the
scan died at the first mismatching socket, so a second client's handshake
ACK never reached its socket and it hung in SYNRCVD - tcpdump showed our
SYN+ACK going out and the client's NICK retransmitted five times into
silence.
- listen() only conflicts with another LISTENER on the port; connections
share it by design now.
- MAX_SOCKS 4 -> 6 (listener + 3 clients + 2 spare; 537 bytes of WRAM0 still
free), and SK_ACC marks a spawned connection as not yet accepted.
tcp_close's settle is also time-boxed to ~250 ms of real time. It counted
PUMPS (8 x 8000), which is ~30s of wall clock when the peer never answers -
invisible with well-behaved peers, but a daemon refusing a client that stayed
connected froze itself, and every other client with it, for half a minute.
httpd keeps its listener open across requests now, so the next client's SYN
is accepted immediately instead of being dropped between close and re-listen:
three parallel clients go 0.1/1.1/2.1s -> 0.4/1.0/0.1s.
|
|
Reported by an outside agent (BUG-schedyield-clobbers-hl.md) while putting
gbos in a browser, where there is no link port at all. Verified here, fixed,
and re-verified.
SchedYield had two exits. The switch path preserved registers because
hSwitchTo saves and restores the full context. The "nobody else is runnable"
path was a bare `ret c` straight out of FindNextReady, which leaves HL
pointing into wProcTable (PcbPtr puts it there). Two callers - net_op_recv's
timeout loop and tcp_connect's - keep a pointer in HL across that call:
ld hl, wNetTO
inc [hl]
jr nz, .wait
call SchedYield ; HL now = &wProcTable, not &wNetTO
inc hl
inc [hl] ; stray increment into a PCB
With no DHCP server the boot-time `dhcp` spins in that loop for ~10s, and
with the shell blocked on tty input nothing else is PS_READY - so *every*
yield took the no-switch path and corrupted a byte, until a saved return
address rolled from $00xx to $01xx and RET landed on the cart entry $0100.
A clean-looking warm reboot, every ~11s, forever. It never happens with
tools/gbhub running, which is why it survived this long.
Fixed in the scheduler rather than at the two call sites, so the contract
holds for every present and future caller: SchedYield now preserves AF/BC/DE
and HL on both paths. On the switch path the pushes sit on the task's own
stack and are popped when it is rescheduled.
Evidence, with the link port unconnected (no --serial-sock; note --headless
bypasses the emulator's breakpoints, so these run on the display path):
before: breakpoint $0100 hit 5x in 40s, first at cycles=46,286,496
(matching the report), deltas ~45.5M cycles apart; HL at the
`inc hl` after the call read $C000 instead of $C9D1
after: 0 hits in 60s, HL reads $C9D1 every time, and a 60s screen watch
shows no spontaneous reboot (it reproduced at 9.9s before)
No regression: 34/34 across the tool, httpd and client suites.
|
|
/# httpd &
host$ curl http://10.0.0.2/ -> a listing of the cart RAM
host$ curl http://10.0.0.2/readme -> the file, from battery-backed SRAM
GET only. "/" serves /index.html if present, else a generated listing; any
directory lists with every entry linked, so you can click around a 1997
handheld's disk from a browser. Content type from the extension. No
Content-Length: HTTP/1.0 + Connection: close delimits the body by closing,
which suits a stack with one in-flight segment.
All output goes through one 200-byte buffer, so each net_send is a full
segment instead of one tiny segment per string.
It drains the whole request, not just the request line. That is not
politeness: leaving the rest unread means those segments are never
acknowledged, so the client retransmits headers for a minute, never sends
its FIN, and its dead connection keeps arriving at port 80 long after we
re-listen. One curl with 1.4 KB of headers stalled every following request
for 30-120s.
It deliberately does NOT read the keyboard either. Backgrounded - the useful
way to run it - every key it polled would be a key stolen from the shell's
prompt. It is a daemon: kill(1) stops it, and the kernel reclaims its socket.
Verified with curl: index, byte-exact file bodies, nested paths, directory
listings, 404, 501, a 1.4 KB request spanning segments, browser user-agents,
three parallel clients (0.1/1.1/2.1s), the shell staying usable alongside it,
and six kill/restart cycles - more than there are sockets.
|
|
The stack could only ever dial out. Two new ops give it an ear:
- listen(sock, port): sockets demux on LOCAL PORT alone, so the listener IS
the connection - the SYN fills in the peer and the socket becomes the
connection. One at a time, re-armed by listening again after close, which
is all a 4-socket table can honestly promise. A second client's SYN is
dropped in silence; its retransmit is taken once we are free (~1s), so
three parallel clients complete in 0.1/1.1/2.1s.
- accept(sock): one RX pump per call, never blocks (like recv_nb), so a
server keeps its own main loop and stays killable.
- TCP_LISTEN/TCP_SYNRCVD in tcp_in: answer a SYN with SYN+ACK, become
established on the handshake's ACK, and process any data that ACK carries.
tcp_peer_is_current tells a peer's retransmitted SYN from a newcomer's.
Without it, "any SYN in SYNRCVD is a retransmission" let a second client
*hijack* a connection mid-request: the first client got its SYN+ACK, sent
its GET twice, and was never answered (found in a tcpdump of two parallel
curls).
listen() also refuses a port another socket already holds - two sockets on
one port silently black-hole the second listener, since demux takes the
first match.
Sockets now record the pid that opened them (SK_OWNER) and exit *and* kill
release them: with 4 slots, a few killed irc/nc/httpd runs used to leave the
box unable to open a socket at all until reboot.
tcp_close stops settling the moment the peer's FIN/ACK arrives instead of
burning its whole pump budget. It otherwise sits in FINWAIT for about a
second, and a server that closes and re-listens drops the next client's SYN
in that window: every request after the first paid a SYN retransmit.
1.1s -> 0.1s per request.
|
|
sys_sleep kept its entire state in three globals (wSleepTarget, wSleepAcc,
wSleepLastDiv) and yielded inside its own loop. With two processes sleeping
at once, each zeroed the other's accumulator on entry, so neither ever
reached its target: both slept ~forever.
Nothing hit it while only one program at a time paced itself. `httpd &`
(polling accept) plus a foreground `ping` is enough: ping wedged with its
echo reply already sitting in the socket buffer, unread.
Now the deadline lives in registers - i.e. on the caller's own stack across
SchedYield, which makes it per-process by construction - and it counts the
64 Hz tick (wTicks, via TimerISR) instead of accumulating DIV deltas.
ticks16 re-reads the counter until the high byte is stable, so a carry
between the two byte reads can't report a 256-unit jump.
|
|
sys_getb/sys_putb indexed inode.blocks[pos/256] with no limit at all, so a
file that grew past its pointers just kept walking: first through the 4
unused bytes at the tail of the 16-byte inode, then straight into the NEXT
inode, reading its type/nlink/size as block numbers.
It hides well. Reads and writes alias identically, so a big file can be
written and read back byte-for-byte and look fine - until something else
touches the neighbouring inode, after which the tail of the file is garbage
from whatever block those bytes now name. Found by serving a 7 KB file over
httpd: corruption began at exactly offset 3072, and only sometimes.
- NDIRECT 8 -> 12: bytes 4..15 are all block pointers now, which is what the
runaway indexing was already doing by accident. Files go to 3 KiB, no
on-disk layout change, no format bump.
- getb past the last direct block reports EOF; putb drops the byte like a
full disk. A capped file beats a corrupted neighbour.
count 250 > big (7142 bytes) now stops at 3072 and its neighbours survive.
|
|
The shell has had pipes for a while with nothing to pipe *through*, and
an 18-row screen with no way to stop output scrolling off it.
- grep [-vin] PAT [file]: substring match (no regex), stdin or a file,
grep(1)'s exit convention. The other end of the pipe, at last.
- more [file]: pages 17 lines at a time, --More-- bar in white-on-blue,
any key pages, q quits. Counts *screen* rows, so wrapped lines pay
their real cost. Keys come from pollin/pollcon, not stdin, so
`cat big | more` still has a keyboard.
- cp SRC DST: the fs has rm/mkdir but no way to duplicate a file. mv is
this plus rm (there is no rename). Rejects SRC == DST, which would
otherwise truncate the source via O_WRITE before reading a byte.
- nc [-s] HOST PORT: raw TCP. Interactive (START sends the line with a
CRLF, /q quits) or -s to pump stdin and print the reply with its own
wall-clock idle timeout, because the kernel's blocking recv counts
pumps, not seconds. Generalizes what chat/irc hardcode:
`nc towel.blinkenlights.nl 23` works.
sys_pollcon now pumps the link before draining the console ring. The
ring is filled *by* net_pump, so a program that never touches the
network polled a ring nothing would ever fill: host-injected keys
(gbtype/gbdemo) hung `more` forever, and the Game Boy stopped answering
pings for as long as it ran. KGetc already pumps for this reason.
New tools take banks 35-38. Locals over statics in all four: a program's
_DATA starts at $A000, where the shell leaves the command line.
|
|
Found by writing nc(1): a 96-byte receive buffer against a host that
answers in 200-byte segments hung the process every time.
- recv/recv_nb copied RXLEN bytes into the caller's buffer and never
looked at the caller's maxlen, so any short buffer got overrun and its
stack smashed. Deliver min(RXLEN, maxlen) and keep the tail in RXBUF
for the next call. wget only survived because its buffer (220) happens
to exceed our MSS.
- tcp_buffer_data overwrote RXBUF from offset 0 on every segment, so a
segment arriving before the app drained the previous one destroyed
unread bytes - silent, undetectable data loss. Append at RXBUF+RXLEN
instead, and drop (without advancing rcv_nxt) when it doesn't fit, so
the peer retransmits once there's room. RXLEN is one byte, so 255 is
the capacity that matters, not RXBUF's 288.
- A draining recv now clears RXLEN as well as HASRX: RXLEN is the append
offset, and a stale one made every later segment look too big to fit,
stalling the connection until the peer gave up.
- A FIN past a gap is not end-of-stream. tcp_in accepted any FIN whose
segment carried no data, without a sequence check, marking the socket
DONE and making recv report EOF at the hole - truncating a transfer
right before its last segments. Decide in-order-ness once, before
rcv_nxt moves, and re-ACK an out-of-order FIN instead.
Verified against a 6300-byte reply received while sending 831 bytes:
byte-exact (6339 = payload + banner, 303 lines, first/mid/last intact),
plus no regression in ping / nslookup / wget / irc.
|
|
sys_write sets wBatch; render_cursor_tile - the single choke point all
of term_putc's visible mutations pass through - then marks a dirty
span (per LINE-SLOT, so spans survive scroll rotation for free)
instead of rendering. term_write_end renders every dirty tile exactly
once, then the cursor. Scrolls stay immediate (blank rebuild is cheap
and artifact-free); file-bound writes flush for free.
The enabler is in libc: puts() was a SYS_PUTC trap PER CHARACTER, so
almost no console output ever went through sys_write. It now issues
one SYS_WRITE for the whole string - one trap, one batch, and every
line-printer (echo, irc, the shell) gets the batched path.
Measured: 37-char write = 128k cycles (3.5k/char) vs ~8k/char down
the per-char trap path. Full demo passes; GIF regenerated.
|
|
KEY1 prepare + STOP at KernelInit (skipped on warm reboot if already
fast). The PPU/LCD keep their normal rate; the timer and DIV double
with the CPU, compensated at the two consumers:
- TimerISR now fires at 128 Hz; a toggle byte keeps wTicks at 64 Hz
(uptime semantics unchanged)
- sys_sleep halves the DIV-delta accumulator's high byte (256 DIV
ticks = 1/128 s in double speed) so gsleep/msleep stay real-time
Measured: count 60 (scroll-heavy) 3.77s -> 1.89s wall; wTicks 64 Hz
over 4s; ping's 800ms msleep pacing unchanged (2.77s for 4 echoes).
sl0pboy already emulated KEY1/STOP + per-speed timer/PPU rates.
|
|
The ring-scroll wrote SCY mid-frame; SCY is sampled per scanline (on
hardware and in sl0pboy's PPU), so a scroll landing mid-frame rendered
the top of the frame at the old offset and the bottom at the new one -
a one-frame shear, invisible in frame-sampled GIF captures but ugly on
a live sixel view.
term_view_update now writes hSCY (HRAM) and a transparent VBlank ISR
(push af / apply / pop af / reti, same profile as TimerISR) copies it
to rSCY, so the view only ever moves at frame boundaries. The scroll's
tile+map writes stay immediate: the new map row is invisible at the
old SCY by construction (it's the ring row one past the visible 18).
|
|
The BG map's 32 rows become a ring: logical row L always lives at map
row (wMapTop + L) & 31, and SCY = (wMapTop + wViewTop)*8 places the
view. Scrolling bumps wMapTop, rebuilds the one new bottom line and
writes its single map row (map_row_one); the other 17 rows move for
free. The whole OSK view dance - shift the cursor above the keys,
restore the backlog on hide - is now term_view_update: one SCY write,
no map rewrite at all.
OSK safety: the keyboard lives on the WINDOW layer, which SCY never
moves. Stale ring rows can only appear under it: wViewTop = max(0,
wCurRow-14) <= 3 is nonzero only while the OSK is visible, and the
window covers exactly those bottom 3 rows (WX=7, WY=120). Verified:
scroll + view shift + scroll-while-OSK-up + backlog restore all
pixel-correct via VRAM screenshots; full README demo passes.
term_scroll: 78.5k -> 35.6k T-cycles (857k pre-optimization: 24x).
OSK toggle: 45k tilemap rewrite -> 212 cycles.
Regenerated demo/gbos-demo.gif on the new renderer.
|
|
1. color_setup memoization: a call with the same (colL,colR) pair as
last time returns immediately (masks/attr still valid in WRAM).
Runs of same-colored cells - i.e. almost all text - hit this.
2. build_tile blank fast path: two spaces = bg-only planes, 8 constant
rows; term_scroll's cleared line and every blank region skip the
glyph pipeline entirely.
3. build_tile two-phase rewrite: combine both glyphs into wRowBuf with
pointers in registers (the old per-row 16-bit WRAM pointer walk was
most of the blit), then compose planes unrolled with plane-0 masks
in B/C.
4. term_write_tilemap attr pass: split each row at the pos-256 VRAM
bank boundary into two tight cache->tilemap copy runs - no per-tile
addressing or bank test.
Measured (emulator cycle counter, ANSI test screen):
term_scroll 857k -> 248k (attr cache) -> 78.5k T-cycles
term_write_tilemap 723k -> 114k -> 45.4k
term_putc 18.2k -> 6.7k
A scroll is now ~1.1 frames of guest CPU (was 12); a full 40-char line
prints in ~63ms of guest time (was ~173ms).
|
|
term_write_tilemap's attribute pass ran color_setup - the full Fano
palette walk - for all 360 tiles on every scroll. render_tile already
computes the attr when a tile changes, so store it (COL_BUF + slot*64
+ ACACHE + tcol; the stride's 24 spare bytes were free) and the attr
pass becomes a table read. update_cursor_attr reads the cache too.
term_init now redraws before writing the tilemap so the cache is warm.
Measured entry-to-return with the emulator cycle counter, ANSI test
screen content: term_write_tilemap 723k -> 114k T-cycles (6.3x),
term_scroll 857k -> 248k (3.5x) - a scroll drops from ~12 frames of
guest CPU to ~3.5.
|
|
Fano-plane palette scheme: the 7 colors map onto 7 CGB BG palettes so
any color pair shares one palette; a tile's palette is picked from the
set of colors its two cells need (tables generated by tools/gencolor.py).
Per cell a packed (bg<<4)|fg byte lives alongside the char shadow, and
the glyph blitter steers glyph/empty pixels to each cell's fg/bg color
slots via plane masks.
An ANSI-ish CSI parser (ESC [ .. m) drives it; usr/ansi.c demos it and
the irc client now renders hashed nick colors, status dimming and a
channel-activity bar.
|
|
The kernel had no clock: TimerISR was a reti stub, IEF_TIMER masked, and
IME was never enabled - the vectors were decorative. sys_sleep just polls
DIV deltas per-process; nothing counted globally.
Now: TAC runs the hardware timer at 16384 Hz with TMA=0, so TIMA overflows
at exactly 64 Hz; TimerISR increments a monotonic 32-bit wTicks (wraps
after ~2.1 years). Scheduling stays cooperative - the ISR is transparent.
Enabling IME in a kernel written for zero interrupts needs care wherever
SP points into memory whose bank is being switched (an IRQ pushes onto SP):
- read_block/write_block map the disk bank over the $A000 window that
holds the caller's stack -> di/ei around the transfer (~2ms, well under
the 15.6ms tick period, so no tick is ever lost)
- hSwitchTo switches SVBK + cart-RAM banks under the outgoing stack ->
di on entry, ei once the incoming stack is mapped
- fork already runs on KSTACK_TOP2 (fixed WRAM) - safe as-is
- term_putc's SVBK switch only remaps $Dxxx, stacks live in $Axxx/$Cxxx
SYS_UPTIME (36) copies the counter (4B LE, di/ei so the read can't tear)
to a user buffer; libc gticks(); usr/uptime.c formats 'up [Nd] H:MM:SS'.
uptime avoids SDCC long div/shift entirely: sm83.lib modules link into
their own areas that land in the $A000 RAM window (latent build.sh trap,
documented there) - bytewise >>6 plus bounded subtraction loops instead.
|
|
Print everything the kernel knows implicitly at boot, Linux-flavored, on
the LCD console and mirrored over the link (gbhub logs each GB's boot):
gbos sm83 microkernel
console: CGB (boot a=$11) <- boot ROM's A/B, saved at entry
cart: mbc5, 1M rom, 128K sram <- our own cart header ($0147-49)
mem: 32K wram 16K vram 127B hram <- CGB constants
proc: 8 slots, 31 programs <- link-time table sizes
net: slip on link port, 4 sockets
fs: gbfs v3 mounted, 120/128 blk free <- bitmap popcount; 'formatted'
tty: 40x18 console, SELECT = osk on first boot
init: spawning pid 1
New src/dmesg.asm with kernel print helpers (kputc/kputs/kputhex/kputdec)
that bypass KPutc (no process context at boot). term_init now runs before
net/fs init so the spew is visible as subsystems come up.
|
|
Userland sources live in usr/ (fits the Unix theme better than 'c').
SDCC output .bin blobs land in build/usr/ with the other build artifacts
instead of littering the source dir; programs.asm INCBINs them from there.
Byte-identical ROM.
|
|
Retransmitted segments (GB ACKs lag behind slow LCD rendering, so real
servers do retransmit) were accepted as fresh data: the same line rendered
again on every retransmit, and rcv_nxt over-advanced so every later outgoing
segment carried an ACK beyond the peer's snd_nxt - which real stacks drop,
silently wedging the session ('/join does nothing' until reconnect).
|
|
|
|
irc HOST [NICK] - connects over the kernel TCP stack (DNS-resolves the
host), registers, and runs a live client on the 40x18 LCD.
UI, within the terminal's means (no cursor addressing - just \r + \b):
messages scroll above a fixed irssi-style input line '[#chan] text_'
redrawn in place; long input scrolls horizontally. The elders' formats:
<nick> msg, <nick:#c> off-channel, *nick* private, -nick- notice,
* nick action, >target< outbound, -!- server/status. Keys come from both
the OSK (SELECT) and the console ring (pollcon), so a hub can drive it.
Commands: /join /part /msg /me /nick /quit /raw, plus bare text to the
current channel. Handles PING (PONG + a wink), CTCP ACTION/VERSION,
JOIN/PART/QUIT/KICK/NICK, 332/353 topic+names, and 433 nick-in-use
(auto-appends _). Registers on bank 32 as program id 30.
Note: gbos doesn't zero C statics, so main() inits its state explicitly;
the local TCP port is randomized (DIV) to dodge a stale server-side
half-open from an unclean prior exit. Tested end to end against a small
ircd through gbhub: full MOTD burst, join, channel + private messages,
actions, and bot replies all render correctly.
|
|
Programs that own their main loop (the IRC client) need to poll for
typed input without blocking. pollin() only sees the on-screen keyboard;
bytes injected over the link (gbhub 'type', for scripted/hub-driven
sessions) land in the kernel console ring, previously only drained by the
blocking KGetc path. Add SYS_POLLCON: a non-blocking con_pop for
userland.
Also enlarge that console ring 16 -> 64. One net_pump drains an entire
serial burst into the ring at once, so a whole injected command line has
to fit or bytes are dropped and lines merge (a 20-char command came out
truncated and glued to the next). 64 covers a full line; mask stays a
power of two.
|
|
SK_STATE/SK_SND/SK_RCV sat at struct offsets 305/306/310, past the
288-byte SK_RXBUF. But the field accessors add the offset as an 8-bit
immediate (add SK_x; ld l,a), so rgbasm truncated 305->49, 306->50,
310->54 (the -Wtruncation warnings we'd been ignoring) - putting STATE
and the sequence numbers *inside* RXBUF at offsets 49/50/54.
Any TCP segment with >=33 payload bytes therefore overwrote the
connection's own STATE and rcv_nxt/snd_nxt with message text. The first
segment of a stream landed, its data clobbered STATE to a garbage value,
and every subsequent segment was dropped because tcp_in no longer saw
ESTABLISHED - so multi-segment TCP receives (an IRC MOTD, any HTTP body
past one segment) silently stalled. wget appeared to 'work' only because
a single-segment reply plus a never-honored FIN still printed once.
Fix: reorder the struct so STATE/SND/RCV precede the big RXBUF, keeping
every field offset < 256 (SK_SIZE unchanged at 314, all accessors are
symbolic). Zero -Wtruncation warnings remain. Verified an 11-line IRC
registration burst now arrives intact.
|
|
Pinging sl0p.foo through gbhub looked hung: DNS resolved, then nothing.
Packet-tracing showed the echo request leaving the hub's TUN and eth0
correctly NATed every time - but 80.78.19.56 blackholes ICMP for all
ids in windows of tens of seconds (provider rate limiting). One lost
reply wedged ping for minutes because net_recv's pump-counted timeout
is effectively unbounded at native emulation speed.
Two fixes:
- NET_RECVNB (op 8): non-blocking recv - one RX pump, $FE if nothing
buffered, same delivery/EOF semantics as NET_RECV otherwise. ping now
waits <=~2s per seq (net_recv_nb + msleep loop), prints 'seq=N
timeout' and moves on, like real ping.
- ICMP echo id was hardcoded $1234 for every GB, every boot, so all
sessions produced byte-identical flows - hostile to NAT conntrack
(keyed on icmp id). wNetEchoId is now our host octet + rDIV timing
noise sampled at DHCP lease, distinct per GB and per boot.
Verified: 8 back-to-back native-speed gbhub runs, zero hangs; a run
that hit a blackhole window printed seq=1 timeout then recovered to
3/4 received. GB<->GB ping and DNS/DHCP unaffected.
|
|
A Game Boy sitting at the shell prompt now services the network instead of
being deaf until you run netd: KGetc (the console input wait) pumps net_pump
each poll, so inbound pings are auto-answered while idle. net_pump gained an
in-frame model and routes any non-framed link bytes to a small console-input
ring (con_push/con_pop), so a headless-injected command still reaches the shell
while SLIP frames go to the stack. Removes the old SLIP-skip-in-KGetc hack.
Verified: two Game Boys on the hub, GB0 idle at the prompt (no netd), GB1
`ping 10.0.0.2` -> 4/4 replies, routed GB1->hub->GB0->hub->GB1.
(Separate, pre-existing: gbhub *spawn* mode and windowed gbjoin can drop an idle
link socket - under investigation; daemon mode + headless is solid.)
|
|
The address is no longer baked in. net_init starts at 0.0.0.0; a DHCP client
runs as the first thing on boot (init/shell forks it and waits), and only once
it has a lease (or gives up) does the prompt appear.
Kernel:
- IP is configurable: 0.0.0.0 until leased; NET_SETIP op stores it.
- 16-bit frame length. DHCP/BOOTP packets are ~272 bytes, over the old 255-byte
frame cap, so net_slip_send takes a 16-bit length, net_pump assembles into a
320-byte buffer with a 16-bit wNetRxLen, and udp_send writes a 16-bit IP total.
Payloads stay <=255 (kept small on purpose) so the per-protocol datalen math
is unchanged. wNetTx/wNetRxBuf 256->320, SK_RXBUF 208->288.
Userland:
- c/dhcp.c: DISCOVER->OFFER->REQUEST->ACK over a UDP socket, then net_setip();
times out gracefully (shell still boots) if there's no server. sh.c runs it
before the prompt.
Bridge (self-contained DHCP server, no dnsmasq):
- tunbridge.py + netboot intercept UDP->:67 and answer OFFER/ACK leasing
10.0.0.2 (gateway 10.0.0.1); everything else is bridged/NATed as before.
Regression-tested ICMP/UDP/TCP after the 16-bit change. Verified end to end:
dhcp: discovering
dhcp: leased 10.0.0.2
/# ping 10.0.0.1 -> 4/4 (traffic from the leased address)
|
|
There was no time source at all (IRQ vectors just reti; scheduler is purely
cooperative). The DIV register (FF04) free-runs at 16384 Hz regardless of
interrupts, so sys_sleep accumulates DIV deltas across SchedYields (other procs
keep running) until the requested number of 1/64-second units elapse.
- SYS_SLEEP(34): B = 1/64s units; libc gsleep(units) / msleep(ms) wrappers.
- ping now msleep(800) between echoes, so it paces like real ping instead of
blasting all four at once.
Verified real-time (capped emulator): replies land ~0.85s apart. In --uncapped
runs the delay is GB-time (fast wall-clock), as expected.
|
|
The link port is both the console-input fallback and the network. At the shell
prompt the host's multicast (mDNS/LLMNR/IGMP) was fed to the serial and KGetc
read those packet bytes as console input -> garbage in the shell, OSK unusable.
Two-part fix:
- KGetc now skips SLIP frames (0xC0-delimited) on the serial console path, so
inbound packets never surface as console input. Plain injected bytes (the
headless tunbridge command path) still pass through. New wConInSkip flag.
- netboot/tunbridge only forward IP packets destined to 10.0.0.2 (drop the
multicast noise at the bridge).
Verified: GB sits at a clean "/#" prompt while all packets (noise included) are
forwarded; being-pinged (3/3) and wget still work. On the emulator, the OSK is
SELECT=space, START=enter, A=z, B=x, d-pad=WASD/arrows.
|
|
Kernel TCP client on the socket layer: net_connect(SOCK_TCP) runs the 3-way
handshake in-kernel, send()/recv() drive the byte stream, close() does the FIN.
- Socket gains state + 32-bit snd_nxt/rcv_nxt (big-endian, add-with-carry-fold).
- tcp_send_seg builds IP+TCP with a pseudo-header checksum; the SYN carries an
MSS option (200) so the peer never sends a segment larger than our 256-byte
frame buffer (we don't do IP reassembly).
- tcp_in state machine: SYN_SENT->ESTABLISHED on SYN-ACK, buffers in-order data
and ACKs it, handles FIN -> recv() returns 0 (EOF).
- No retransmission: the GB<->host link is lossless and the host's real TCP
owns the internet side - which removes TCP's hardest part.
- net_pump now processes one frame per call so recv drains each segment before
the next arrives (single rx slot, no overwrite).
wget.c is now a thin socket client: connect -> send "GET / HTTP/1.0" -> recv to
EOF -> print. Verified end to end against a host HTTP server:
/# wget 10.0.0.1
HTTP/1.0 200 OK
Hello from a real HTTP server, fetched by a Game Boy!
with a clean SYN/SYN-ACK/ACK ... PSH ... FIN/ACK trace on the wire.
|
|
Adds UDP to the kernel socket layer on top of the ICMP core:
- net_sum() (raw folded sum) split out of net_cksum() so a UDP pseudo-header
(src/dst IP + proto + length) can seed the segment checksum.
- udp_send: builds IP+UDP with the pseudo-header checksum; NET_BIND sets the
local/source port; net_connect sets the peer.
- udp_in: demuxes inbound UDP by destination port to the bound socket
(net_find_udp), delivers the payload + source addr.
Also fixes a real recv bug: net_pump clobbers BC/DE/HL, so the old recv timeout
counted in registers and was effectively random. recv now counts in WRAM.
New `nslookup HOST` (PROG_NSLOOKUP=28, bank 30): builds a DNS A query and parses
the answer (with 0xC0 name-compression) entirely in userland over a UDP socket -
the kernel never sees DNS, just UDP. Verified through the bridge NAT:
/# nslookup example.com -> example.com -> 172.66.147.243
This gives us name resolution for the TCP/HTTP demo next.
|
|
The network stack moves into the kernel. src/socket.asm owns SLIP framing,
IPv4, RFC1071 checksums, and ICMP; programs now speak a socket API through one
syscall (SYS_NET, DE=&netreq dispatched by op): net_socket/connect/send/recv/
close/poll (c/sock.h). No program touches SLIP, IP headers, or checksums.
- Socket table (4 sockets) + tx/rx buffers in WRAM0; our IP = 10.0.0.2.
- net_pump: drains the link, reassembles SLIP frames, demuxes IPv4. Inbound
ICMP echo requests are auto-answered in-kernel, so the GB replies to pings
whenever any process pumps RX.
- ICMP sockets: send() emits an echo request to the connected peer; recv()
returns the matching reply (with a spin/yield timeout).
ping.c is now a ~15-line socket client; netd.c is just `for(;;){net_poll();
yield();}`. Verified over tunbridge:
/# ping 1.1.1.1 -> replies from the real internet (kernel builds it all)
host# ping 10.0.0.2 -> 4/4, 0% loss (kernel auto-answers)
Gotchas recorded: gbos.inc isn't a make dep (touch asm after editing); this
crt0 doesn't copy initializers (fill arrays at runtime); and the arg string at
0xA000 overlaps _DATA, so parse targets must sit past it (big buffer first).
UDP and TCP sockets build on this same core next.
|
|
New `ping [A.B.C.D]` program (PROG_PING=27, bank 29): builds and sends ICMP
echo requests from 10.0.0.2, then reads replies off the link port. It keeps
reading SLIP frames until it finds *our* echo reply, skipping the IGMP/mDNS/
SSDP multicast noise that shares 10.0.0.0/24. Reply wait uses a generous
srecv_nb spin budget since the emulator runs uncapped (no timer syscall yet).
Verified over the tunbridge (with NAT):
/# ping 10.0.0.1 -> 4/4 received, ttl=64 (the SLIP peer/host)
/# ping 1.1.1.1 -> 4/4 received, ttl=56 (Cloudflare, real net!)
ttl=56 is a real internet round trip (64 minus the hops). Combined with the
host being able to ping the GB, the Game Boy is now a full two-way ICMP host.
|
|
Start of an actual TCP/IP stack on gbos (TLS stays in a proxy). netd is a
userland IP responder over SLIP: our address is 10.0.0.2, the SLIP peer
10.0.0.1. It parses IPv4 headers, answers ICMP echo requests, and rebuilds
the packet with correct IP + ICMP checksums (RFC 1071 one's-complement sum,
carry-folded - works fine on the SM83).
c/netd.c + register; tools/gateway.py gains --mode ping: it crafts ICMP echo
requests over SLIP and verifies the replies.
Verified: `netd` answers 4 pings, gateway reports reply from 10.0.0.2 with
cksum=ok for each. Next: UDP, then TCP.
|
|
The application layer of the link-port demo, and it ties the whole system
together: the LCD terminal displays, the on-screen keyboard types, and the
link port carries a live chat.
Kernel: sys_srecv_nb (non-blocking link receive; A=byte, CF=none) and
sys_pollin (poll the OSK for a typed char without blocking) - syscalls
31/32. Both are what a poll loop needs to receive and type at once.
Userland: c/chat.c runs a poll loop - it feeds non-blocking bytes through a
SLIP receive state machine and prints whole incoming frames as messages,
while pollin() drives the on-screen keyboard; SELECT shows the keys, type a
line, START sends it as a frame. libc srecv_nb()/pollin().
Host: tools/gateway.py --mode chat is a simple bot peer (echoes each GB
message and injects a few async ones); --keys can drive the OSK for tests.
Verified: the gateway pushes 'welcome', '<alice> hey gameboy!', '<bob> nice
link cable' unprompted and the GB displays all three (async receive); typing
'hi' on the OSK echoes it and emits the SLIP frame \xC0hi\xC0 (send). A Game
Boy in the chat, keyboard on screen, over the link cable.
|
|
Grow the link-port demo from echo to actual network access. The Game Boy
still only does SLIP framing + display; the host gateway does DNS/TCP/HTTP.
- c/netlib.h: SLIP framing factored out (header-only, per-program copy).
necho.c now uses it too.
- c/wget.c: `wget URL` frames the URL, then prints the reply body. The
gateway streams the body back as typed frames: 'D'<chunk> ... 'E'.
- tools/gateway.py: add --mode http (urlopen the frame as a URL, cap the
body, chunk it) alongside --mode echo; --cmd runs any gbos command.
Verified: `wget example.com` streams back the full Example Domain HTML onto
the terminal; `wget sl0p.foo` fetches the real page. A Game Boy on the web,
over the link cable.
|
|
First step of link-port networking. The LCD terminal + OSK freed the serial
port from console duty, so it can be the network link.
Kernel (src/net.asm): raw link-port serial that bypasses the console/
terminal - sys_ssend (transmit, GB drives the clock) and sys_srecv (receive,
GB slave, blocks by yielding). Syscalls 29/30.
Userland: libc ssend()/srecv(); c/necho.c does SLIP (RFC 1055) framing over
them - send a packet, receive the reply, print it.
Host: tools/gateway.py wraps the emulator, owns its link serial, speaks SLIP,
and (for now) echoes every frame back - the "link cable adapter". Console
(ASCII) bytes on the same channel are printed for visibility.
Verified: `necho` sends a SLIP frame, the gateway decodes+echoes it, and gbos
prints the reply - a real framed round-trip over the Game Boy link port.
Next: swap the echo for actual network ops (DNS/HTTP or IRC/chat).
|
|
Bump PIPE_MAX to 4 so a|b|c|d|e (4 pipes) works. This stays within the
single-byte buffer-offset math (idx*64+pos <= 3*64+63 = 255) and the fd
space ($F0..$F7, clear of $FF console).
The bump exposed a data-corruption bug that also affected the 2-pipe case
(just invisibly - a wc-only test can't see mangled bytes): pipe_write kept
the byte-to-write in the SHARED wPipeByte global across its SchedYield
(buffer full), so a concurrent pipe op clobbered it and the writer then
stored the wrong byte. Now the byte is held in D across the yield, and
pipe_bufptr no longer clobbers D; pipe_read/pipe_write also push their idx
across SchedYield rather than assume the yield preserves registers.
Verified: count N | cat now streams EXACT content (no 'linn'/'llne'
corruption); count 60 | cat | cat | wc = 60 180 1671; 5-stage
count 4 | cat | cat | cat | cat prints line 1..4; SIGPIPE (count 200|true)
and count 100|wc still fine, no hangs.
|
|
Replace the shell's temp-file pipe hack with proper in-kernel FIFOs.
Kernel (src/pipe.asm, new):
- A small pool of bounded ring buffers (PIPE_MAX=2, 64 B each) with
writers/readers refcounts. Pipe fds are $F0+idx*2 (read) / +1 (write);
$FF stays "console".
- SYS_PIPE allocates one (writers=readers=1) and returns the read fd
(write = read+1). getb/putb/close dispatch pipe fds here; sys_exit drops
the refcounts held as PROC_STDIN/PROC_STDOUT.
- Blocking with SchedYield, which is the flow control: read blocks while
empty with a writer (EOF once writers hit 0), write blocks while full
with a reader, and if the last reader is gone the writer is killed
(SIGPIPE -> exit 141). Cooperative-scheduler friendly.
Shell (c/sh.c):
- run_pipeline(): split on '|', make a pipe between adjacent stages, and
fork ALL stages concurrently (no wait between), wiring stdin/stdout;
then wait for all. Per-stage >/< still honored; orphaned pipe ends are
closed on a lookup miss so EOF/EPIPE propagate.
- Drop the __pipe temp file and its 2 KB / serialized limits.
libc: pipe(). New c/ptest.c exercises the FIFO (write, read back, EOF).
Now works (old version couldn't): multi-stage a|b|c; streaming beyond 2 KB
(count 120 | wc = 3372 bytes through a 64 B buffer); early-exit SIGPIPE
(count 200 | true kills count instead of hanging/overflowing a file).
|
|
A fixed offset of OSK_ROWS pushed the cursor off the top of the screen when
there was little backlog (e.g. a fresh terminal: cursor at logical row 0 ->
screen row -3, hidden). Now the offset is max(0, wCurRow - 14): 0 when the
cursor is within the visible area, growing only enough to keep the cursor
line at the bottom visible row (14) once the terminal has filled past it.
compute_offset (from wOskVisible + wCurRow) runs at the top of
term_write_tilemap; cursor_down re-runs the tilemap when the cursor changes
rows while the OSK is up; osk_toggle just re-runs it. Shared OSK_DOCK
constant in gbos.inc keeps term.asm and osk.asm in sync.
Verified: fresh terminal + OSK shows the input at row 0; filling past row 14
shifts the view to hold the input at row 14; full screen + OSK puts it at
row 14; toggle on+off is still lossless.
|
|
Replace the shrink-and-scroll region with a view offset. The terminal is
always 18 logical rows; term_write_tilemap maps screen row -> logical row +
wViewTop (rows the OSK covers clamp to the cursor line). Showing the OSK
sets the offset to OSK_ROWS so the cursor line lands at row 14 (above the
keyboard) and hiding it sets the offset back to 0 - now purely a remap, so
nothing is scrolled off or destroyed. term_set_view replaces term_set_rows;
term_scroll/cursor_down are back to plain 18-row scrolling.
Result: full 18-row backlog when the OSK is hidden, a shifted 15-row window
when shown, and toggling is lossless (verified: OSK-shown view == baseline
rows 3..17, and toggle on+off == baseline exactly).
|