Commit Graph
47140 Commits
Author SHA1 Message Date
Henrik Rydgård 25b9b60694 wsdbg: tokenize REPL lines with quote-awareness instead of splitting on any whitespace
handle_repl_line() used line.split_whitespace(), so any parameter value
containing a space (most commonly a logFormat string with multiple
{expression} placeholders, e.g. logFormat="v0={v0} a0={a0}") got silently
split into multiple bogus key=value tokens - build_event_json then either
failed outright ("not in key=value form") or, worse, sent a malformed
request with the value truncated at the first space, with no indication
to the user that their command wasn't parsed as intended. Cost real
debugging time this session before being traced to this (see
docs/VSHBootInvestigation.md).

Added split_shell_words(): a minimal shell-like tokenizer that treats
'single' or "double" quotes as protecting spaces (and the other quote
character) from being treated as a separator, stripping the quotes from
the resulting token; backslash escapes the next character, including
inside quotes. Not a full shell-parsing crate, just enough for this
tool's key=value parameters.

Verified live against PPSSPPHeadless: logFormat="job hit pc={pc} ra={ra}"
now arrives as a single correctly-quoted JSON string value instead of
being split apart.
2026-08-14 11:04:32 +02:00
Henrik Rydgård 8160a7e3eb wsdbg: add --sync to wait for responses instead of guessing sleeps
Scripted sequences piped into the REPL previously had no way to know
when a command's response had arrived, so every multi-step script in
this repo's investigation docs needed "sleep N" between commands,
guessed per case and often wrong in either direction.

--sync makes each line block until its response arrives before the
next line is read: the matching ticketed response for most events,
or (since cpu.resume/stepInto/stepOver/stepOut/runUntil/nextHLE are
all documented as having no immediate response at all - only the
eventual unticketed cpu.stepping broadcast) that broadcast for the
resume/step family specifically. Everything still prints as it
arrives; only the timing of the next send changes. Bounded by
--sync-timeout (default 30s) so a breakpoint that never trips can't
hang a script forever.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01GZq8ZtJmFY7bkX5FVkr3P9
(cherry picked from commit 9ee89fd16ba4557a245ca2a96c69a9c9478cf59e)
2026-08-14 11:04:32 +02:00
Henrik RydgårdandClaude Sonnet 5 2ee4f2fadb Debugger: add gpu.displaylist.disasm - a GE display list decoder over the WebSocket API
Decoding a GE display list previously meant memory.read-ing the raw bytes
and hand-decoding each 32-bit command word against GPU/ge_constants.h's
GECommand enum - which is exactly what it took to find this session's
actual headline VSH boot finding (a display list that clears the screen
once, sets up per-icon render state 6 times, and never issues a single
further draw call - see docs/VSHBootInvestigation.md Attempt 22). That
manual process is real, repeatable, and error-prone by hand; PPSSPP
already has a proper GE disassembler (GPU/GeDisasm.cpp's
GeDisassembleOp(), and GPUCommon::DisassembleOpRange() built on top of
it) used by the ImGui/Windows GE debugger UI - it just wasn't reachable
from the WebSocket API.

New Core/Debugger/WebSocket/GPUDisasmSubscriber.cpp exposes
gpu->DisassembleOpRange() as gpu.displaylist.disasm, mirroring
memory.disasm's own parameter conventions (address+count or
address+end, capped at 10000 commands) and compact mode (one string per
command, "AAAAAAAA  desc", instead of the full {address,cmd,op,desc}
object) added in the previous commit. GE command words live in normal
guest RAM like CPU code, so - unlike gpu.buffer.* - this doesn't require
the CPU/GPU to be paused first, matching memory.disasm's own live-read
behavior.

Added to all 6 build systems that compile the WebSocket debugger
(CMakeLists.txt, Core.vcxproj(.filters), UWP's CoreUWP.vcxproj(.filters),
android/jni/Android.mk - libretro doesn't build any Debugger/WebSocket
files at all, so nothing to add there).

Verified live via PPSSPPHeadless + wsdbg against a real demo ELF: both
compact and full-JSON modes correctly decode real GE command words (NOP/
NOP_FF) with no errors. UnitTest.exe all: 49/49 passed.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01GZq8ZtJmFY7bkX5FVkr3P9
2026-08-14 11:04:32 +02:00
Henrik RydgårdandClaude Sonnet 5 a502b55fea Debugger: add memory.disasm compact mode and memory.searchDisasm findAll
Two real gaps hit repeatedly while investigating the VSH boot path (see
docs/VSHBootInvestigation.md):

- memory.disasm's response is the full per-field JSON (type, address,
  addressSize, encoding, macroEncoding, backgroundColor, name, params,
  symbol, function, dataSymbol, breakpoint, isCurrentPC, branch,
  relevantData, conditionMet, dataAccess - ~15 fields per line). Reading
  disassembly by hand meant writing a throwaway script each time to reduce
  this down to "ADDR: name params" - and at least once, a bug in one of
  those scripts produced misleading output that wasn't caught immediately.
  Added compact=true: returns "lines" as an array of plain strings
  ("M AAAAAAAA  [symbol: ]name params", M = '>' for current PC, '*'/'o'
  for an enabled/disabled breakpoint) instead, computed once correctly
  here instead of ad hoc every time.

- memory.searchDisasm already existed but only ever returned the first
  match - genuinely limiting for "find every caller of this address"
  call-graph-style queries, which came up directly while trying to trace
  which function builds VSH's GE display list. Added findAll=true: scans
  the whole range and returns every match in a new "addresses" array
  (capped at 1000), instead of stopping at the first. Default behavior
  (address: first match or null) is unchanged for existing callers.

Verified live via PPSSPPHeadless + wsdbg: compact mode against a real
demo ELF's entry point produces clean, correctly-marked text lines;
findAll=true against the same range found all 11 jal instructions instead
of just the first. UnitTest.exe all: 49/49 passed.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01GZq8ZtJmFY7bkX5FVkr3P9
2026-08-14 11:04:32 +02:00
Henrik RydgårdandClaude Sonnet 5 be236d47c8 Debugger: surface Core_RequestCPUStep() failure on cpu.stepInto instead of silently hanging
Into()'s same-thread branch called Core_RequestCPUStep(CPUStepType::Into, 1)
without checking its return value. Core_RequestCPUStep() can genuinely
fail (a step/run request is already queued this host frame - see its own
"Can't submit two steps in one host frame" ERROR_LOG) - on failure, no
step happens and no cpu.stepping event ever fires, but cpu.stepInto's own
contract is "no immediate response, a cpu.stepping event follows", so a
rejected request looked identical to a request still in flight: nothing
to distinguish "wait longer" from "this silently failed, nothing is ever
coming." This is part of the same failure family as the delay-slot race
just fixed in PrepareResume() (previous commit) - Core_RequestCPUStep()'s
one-at-a-time guard rejecting a step no caller in this file checked for.

Now calls req.Fail() on rejection so the client gets an explicit answer
instead of an indefinite wait. Updated the cpu.stepInto doc comment to
note the new (retryable) failure mode.

Verified via UnitTest.exe all (49/49).

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01GZq8ZtJmFY7bkX5FVkr3P9
2026-08-14 11:04:32 +02:00
Henrik RydgårdandClaude Sonnet 5 0c63905095 Debugger: report dropped messages when the log broadcaster's ring buffer overflows
DebuggerLogListener buffers up to 1024 log messages between polls of the
WebSocket event loop (up to 1000Hz under high activity, 60Hz otherwise -
see WebSocket.cpp). A source that logs faster than that - a log-only
breakpoint hit thousands of times in a tight loop is a real example, not
hypothetical, see docs/VSHBootInvestigation.md's Attempt 22/24 - can wrap
the ring buffer before GetMessages() ever reads the oldest entries,
silently losing them. From the client's side this was indistinguishable
from the breakpoint just not firing at all, which cost real debugging time
this session tracking down a red herring before finding the real
mechanism.

GetMessages() already detected the overflow case internally (the
`read_ + BUFFER_SIZE < count_` branch) to avoid returning garbage, but
never reported how many messages were actually lost. Now synthesizes a
warning LogMessage ("N log message(s) dropped - client polling too slow
for this volume") and prepends it to the batch whenever this happens, so
a real gap is visibly distinguishable from "this just never got logged."

Verified via UnitTest.exe all (49/49) and a live PPSSPPHeadless + wsdbg
session confirming normal (non-overflow) log relay still works
end-to-end.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01GZq8ZtJmFY7bkX5FVkr3P9
2026-08-14 11:04:32 +02:00
Henrik RydgårdandClaude Sonnet 5 3edea91c8e Debugger: track hit counts on address breakpoints, expose via cpu.breakpoint.list
BreakPoint (cpu.breakpoint.*) had no hit-count tracking at all, unlike
MemCheck (memory.breakpoint.*), which already tracks numHits. This made it
genuinely hard to tell "this breakpoint is never being reached" apart from
"it's being reached but I'm not seeing the log/pause where I'm looking" -
directly informed by repeatedly hitting exactly that ambiguity while
debugging the VSH boot path this session (see docs/VSHBootInvestigation.md).

Added BreakPoint::numHits, incremented in BreakpointManager::ExecBreakPoint()
whenever a breakpoint's address is hit and any condition passes (matching
MemCheck::Apply()'s existing semantics - counts real triggers, not just
"execution passed through here"). Exposed as a new "hits" field in
cpu.breakpoint.list's response.

Verified live via PPSSPPHeadless + wsdbg: hits reads 0 before the CPU
resumes, 1 after the breakpoint fires once. UnitTest.exe all: 49/49 passed.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01GZq8ZtJmFY7bkX5FVkr3P9
2026-08-14 11:04:32 +02:00
Henrik RydgårdandClaude Sonnet 5 4389f706b9 Debugger: fix a real race in cpu.stepOut/stepOver/runUntil/nextHLE from a delay slot
PrepareResume() used Core_RequestCPUStep(CPUStepType::Into, 1) to step past a
delay slot instruction before deciding whether to add a breakpoint and call
Core_Resume() - but Core_RequestCPUStep() only queues that step for
Core_ProcessStepping() to perform later (on the next iteration of the normal
stepping-mode loop). Every caller (Into's cross-thread branch, Over, Out,
RunUntil, HLE) immediately inspected currentMIPS->pc/inDelaySlot right after
PrepareResume() returned to decide what to do next - reading stale,
pre-step state, since the queued step hadn't run yet.

Worse: those callers then call Core_Resume(), which sets coreState back to
CORE_RUNNING_CPU. Core_ProcessStepping() only processes g_cpuStepCommand
when coreState is CORE_STEPPING_CPU/STEPPING_GE/RUNNING_GE, so once resumed,
the queued step is never processed at all - not just late, silently dropped,
leaving g_cpuStepCommand permanently set until the next Core_Break() resets
it. Any cpu.step*/cpu.runUntil request a client issues in that window (CPU
resumed running, breakpoint not yet hit again) hits
Core_RequestCPUStep()'s "Can't submit two steps in one host frame" guard and
is silently ignored, since none of these call sites check its return value -
this is the "step-out sometimes just doesn't do anything" flakiness reported
against this file.

PrepareResume() is only ever called from within a Core_RunOnCPUThread()
callback, so it's always already running on the CPU thread - safe to
single-step synchronously (currentMIPS->SingleStep(), matching how
Core_PerformCPUStep()'s own CPUStepType::Into case does it) instead of
queuing an async request whose completion every caller then assumes without
verifying.

Verified via UnitTest.exe all (49/49). Attempted to force a live repro via
wsdbg against a delay-slot jal in a demo ELF; wasn't able to reliably
trigger the failure window externally (by the time a client's next command
arrives, the CPU has typically already reached its next breakpoint and
Core_Break() has cleaned up the stale state first) - the race window is
real per the code trace above but appears to be narrow enough that it
mainly shows up under real usage timing (a slow-to-reach next breakpoint,
or a fast follow-up command from a script/UI), not simple synchronous
scripting. The fix is unconditionally more correct regardless: it replaces
a fire-and-forget async request every caller immediately assumed had
already completed with a direct synchronous call that actually has by the
time the next line runs.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01GZq8ZtJmFY7bkX5FVkr3P9
2026-08-14 11:04:32 +02:00
Henrik Rydgård b507c9c797 Fix log-only cpu.breakpoint entries always pausing execution
The interpreter's hot-path breakpoint check in
RunUntilDowncountZeroWithChecks called Core_Break() unconditionally
whenever IsAddressBreakPoint() was true - true for any non-ignored
breakpoint, log-only included - instead of routing through
BreakpointManager::ExecBreakPoint(), which is what actually respects
BREAK_ACTION_LOG vs BREAK_ACTION_PAUSE. So a cpu.breakpoint.add with
log=true and enabled=false still paused on hit, contradicting its own
documented behavior.

The JIT backends and IR interpreter don't have this bug - they already
route through ExecBreakPoint() via JitBreakpoint()/IRRunBreakpoint()
and check the result for BREAK_ACTION_PAUSE. Only this one plain
interpreter loop had its own unconditional inline check instead.

Verified: a log-only breakpoint now logs without pausing (325 hits
logged, then execution continued past it normally); a normal enabled
breakpoint still pauses; cpu.stepOver (which relies on temporary
breakpoints, still correctly removed only when they actually pause)
still steps over calls correctly; all 49 unit tests pass.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01GZq8ZtJmFY7bkX5FVkr3P9
(cherry picked from commit 726db5e4ea3651db0eb2de13c622c033dcc95699)
2026-08-14 09:27:21 +02:00
Henrik Rydgård 0c7f8ca07c Poison unknown MMIO reads, make COP0 Count free-run; rule out both for the break
Two diagnostic experiments while chasing the divide-by-zero break from the
previous commits (see docs/VSHBootInvestigation.md "Attempt 8"/"Attempt 9"):

- Unregistered MMIO reads now return a distinctive poison value
  (0x1337BEEF) instead of 0, so a future trace can tell at a glance when a
  value traces back to an unimplemented register instead of looking like an
  ordinary zero.
- COP0 register 9 (Count) now returns a live CoreTiming-derived value
  instead of a static shadow-array read, matching how real hardware free-
  runs it regardless of software writes.

Neither change altered the reboot.bin free-run's outcome at all (identical
break, same PC) - ruling out both as the source of the zero divisor traced
in the previous commit. Kept anyway: both are straightforwardly more
correct/useful than what was there before, independent of this specific
bug.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01GZq8ZtJmFY7bkX5FVkr3P9
(cherry picked from commit ac446449d9031b628d7d9c4dbc101638840f8434)
2026-08-14 09:27:21 +02:00
Henrik RydgårdandClaude Sonnet 5 db2d248b4a Rename GPRBreakpoint/gprBreakpoint to RegBreakpoint/regBreakpoint
The struct and its API only handle GPR indices today, but the naming
should stay general since this is expected to grow to cover other
register files too (e.g. FPU registers like $f10). Pure rename - no
behavior change:

- Core/Debugger/Breakpoints.{h,cpp}: RegBreakpoint struct, all
  BreakpointManager Add/Remove/Change/Get/Exec/Has/Find*RegBreakpoint*
  methods, regBreakpoints_/regBreakpointMask_ members.
- Core/Core.{h,cpp}: BreakReason::RegBreakpoint, "cpu.regBreakpoint"
  break-reason string.
- Core/Debugger/WebSocket/BreakpointSubscriber.{h,cpp}: WebSocket
  events cpu.gprBreakpoint.* -> cpu.regBreakpoint.*, matching
  Add/Update/Remove/List handlers and params struct.
- Core/MIPS/MIPSTables.cpp: local variable names in the interpreter's
  per-instruction breakpoint check.
- docs/WebSocketDebugger.md updated to match.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01GZq8ZtJmFY7bkX5FVkr3P9
2026-08-13 16:09:51 +02:00
Henrik RydgårdandClaude Sonnet 5 34d0417897 Memory breakpoints: normalize the kernel address bit too, not just uncached
Breakpoints.cpp's memcheck-matching NotCached() helpers (used by both the
interpreter's real-time FindMemCheckInRange and the JIT's precomputed
UpdateCachedMemCheckRanges/GetMemCheckRanges) only ever normalized away the
uncached bit (0x40000000), never the kernel bit (0x80000000) - so a
memcheck registered on one kernel/user address alias silently didn't match
a write made through the other. This is a real, general bug (any kernel
code writing through the 0x88xxxxxx-style mirror could dodge a memcheck
set on the corresponding 0x08xxxxxx address), not specific to any one
investigation.

Extended NotCached(u32) to also strip the kernel bit, and added a
NotKernel(MemCheck) counterpart so UpdateCachedMemCheckRanges now expands
each non-VRAM memcheck into all four kernel/uncached combinations instead
of two. VRAM intentionally excluded, matching IsValidAddress's existing
"no kernel-flagged VRAM" comment. cpu.breakpoint (PC) breakpoints are
unchanged - out of scope here.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01GZq8ZtJmFY7bkX5FVkr3P9
2026-08-13 11:37:01 +02:00
Henrik RydgårdandClaude Sonnet 5 28cbb3a967 Add the scratchpad's missing kernel-mode mirrors, fix a mask bug hiding them
The scratchpad (PSP's repurposed-cache scratch RAM, 0x00010000+) was only
mirrored for user-mode access (cached 0x00010000, uncached 0x40010000) -
kernel-mode code sees it at 0x80010000/0xC0010000 (the kernel bit,
0x80000000, is independent of and combinable with the uncached bit,
0x40000000 - not "the uncached bit" as an earlier doc note in this branch
mistakenly called it). Missing entirely from MemMap.cpp's views[] table,
causing a real SIGSEGV the first time kernel-mode code (flash0:/reboot.bin)
touched it.

Adding the two missing views wasn't sufficient: the scratchpad range check
is duplicated eight times (IsValidAddress/IsValid2AlignedAddress/
IsValid4AlignedAddress/MaxSizeAtAddress in MemMap.h, and four more in
MemMapFunctions.cpp), and all eight used a mask (0xBFFFC000) that cleared
the uncached bit but kept the kernel bit, rejecting 0x80010000 as invalid
before ever reaching the now-mapped memory. Fixed all eight to 0x3FFFC000,
matching Memory::MEMVIEW32_MASK (which the JIT backends already used
correctly for the equivalent runtime check).

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01GZq8ZtJmFY7bkX5FVkr3P9
2026-08-13 11:36:42 +02:00
Henrik RydgårdandClaude Sonnet 5 75174af77b Add GPR write breakpoints (break when a register is written, anywhere)
New debugging primitive: break whenever any instruction writes to a
given general-purpose register (0-31), regardless of which address
executes the write. Requested for continuing the reboot.bin trace,
where the actual blocker is "what sets $s3 to this bad value", not
"what happens at a specific address" - existing address/memory
breakpoints can't express that directly.

- GPRBreakpoint (Core/Debugger/Breakpoints.h) mirrors the existing
  BreakPoint/MemCheck shape (result/condition/logFormat/hit count),
  keyed by register index instead of address/range.
- BreakpointManager keeps a u32 bitmask (bit i = register i has an
  active breakpoint) alongside the GPRBreakpoint vector, so the
  interpreter loop can test "would this write trip anything" with a
  single shift+and against a value already cached in a local.
- RunUntilDowncountZeroWithChecks (Core/MIPS/MIPSTables.cpp) computes
  the about-to-be-written register from the current instruction's
  OUT_RT/OUT_RD/OUT_RA flags (GetGPRWriteTarget()) and checks it
  against the mask, same convention as the existing memcheck handling
  right above it (checked before the instruction executes, bails via
  CORE_STEPPING_CPU without running it if tripped).
- New BreakReason::GPRBreakpoint ("cpu.gprBreakpoint") for Core_Break.
- WebSocket API: cpu.gprBreakpoint.add/update/remove/list, accepting
  either a 0-31 'register' index or a case-insensitive 'name' (e.g.
  "s3"), documented in docs/WebSocketDebugger.md.

Interpreter-only for now, deliberately.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01GZq8ZtJmFY7bkX5FVkr3P9
2026-08-13 11:35:29 +02:00
Henrik Rydgård 3f6743d40d Improve Core_ExecException diagnostics per exception type 2026-08-13 11:35:06 +02:00
Henrik RydgårdandClaude Sonnet 5 5dbe9bea91 Add dummy interpreter implementations for the basic COP0 instructions
mfc0/mtc0/rdpgpr/mfmc0/wrpgpr had no interpreter execution function at
all - ordinary PSP user-mode code never executes COP0 instructions directly
so nobody needed one. Real kernel-mode boot code does the opposite:
flash0:/reboot.bin's very first instruction is mfc0.

Int_Cop0 (Interpreter.cpp) backs these with a small file-scope shadow
register array - not real COP0 semantics (no interrupts/exceptions/
TLB), just enough to not fault and give plausible
write-then-read-back-same-value behavior for boot code that pokes
Status and other registers. Might later be added to MIPSState.

Also regenerated Core/MIPS/InterpreterDispatch.cpp (`PPSSPPHeadless
--generate-interpreter-dispatch`).

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01GZq8ZtJmFY7bkX5FVkr3P9
2026-08-13 11:34:00 +02:00
Henrik Rydgård defea54cd2 Merge pull request #22090 from hrydgard/more-stuff
Simplify file system mounting slightly, more MIPS context plumbing
2026-08-13 08:34:27 +02:00
Henrik RydgårdandClaude Sonnet 5 6eb204d882 unittest: support running multiple named tests in one invocation
main() only ever read argv[1], silently ignoring any further test
names despite both the usage text and the file's own top-of-file
comment claiming "one or more of the below" was supported. Rewrote
argument handling to collect every named test (or expand "all"),
validate all names up front (bailing with the usage listing if any is
unrecognized, rather than silently running a partial set), and run
them with the same pass/fail summary "all" already used - single-test
invocations now also get the "**** Running test X ****" banner "all"
always had, for consistency.

Verified: `UnitTest.exe CLZ MathUtil Path` now runs all three (was:
silently only CLZ); a bad name reports "Unknown test: X" plus usage
and exits 1; `all` and single-name behavior otherwise unchanged.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01GZq8ZtJmFY7bkX5FVkr3P9
2026-08-13 08:09:42 +02:00
Henrik Rydgård eb0813c0e3 Add a utility function for all the ABIs to call functions with a pointer arg. Use to call Advance from the JIT with the MIPSContext. Indent some code better. 2026-08-13 08:09:30 +02:00
Henrik Rydgård 95f1dc648d Simplification: Mount filesystems earlier in init, and don't remount on LoadExec
This is needed for an upcoming PR.
2026-08-12 23:38:48 +02:00
Henrik Rydgård 06d5847da3 Merge pull request #22089 from hrydgard/interpreter-mips-state
Interpreter: Plumb through MIPSState *mips as a function parameter, instead of global currentMIPS
2026-08-12 15:47:43 +02:00
Henrik Rydgård 5a8e24f583 Interpreter: Hook up dummy MMIO handlers for future experiments 2026-08-12 14:48:23 +02:00
Henrik Rydgård 6836a2bd4e Libretro buildfix, add a comment 2026-08-12 14:26:10 +02:00
Henrik Rydgård c5e4d0d90d Rename the get-memory-pointer functions to make it clear where CPU exceptions can happen. 2026-08-12 14:06:16 +02:00
Henrik Rydgård e9a3449ede More MIPSState * plumbing (manual) 2026-08-12 14:02:19 +02:00
Henrik Rydgård c9c886d78f Plumb a MIPS context into ReadVector etc. 2026-08-12 14:02:19 +02:00
Henrik Rydgård b10e4ad1b5 Code style improvement, handle unknown instructions with a pseudo CPU exception 2026-08-12 14:02:19 +02:00
Henrik RydgårdandClaude Sonnet 5 ffad0acea1 Plumb an explicit MIPSState *mips through the interpreter's Int_* handlers
All ~82 MIPSInt::Int_* functions (Interpreter.cpp/.h,
InterpreterVFPU.cpp/.h) now take an explicit MIPSState *mips instead
of reaching for the global currentMIPS internally, along with their
file-local helpers (DelayBranchTo, SkipLikely, ApplySwizzleS/T,
ApplyPrefixD/ST, RetainInvalidSwizzleST, EatPrefixes). MIPSInterpretFunc,
Interpret(), ExecInstruction()/InterpreterDispatch.cpp (regenerated),
and RunUntilFast() all thread mips through accordingly.

Deliberately left on currentMIPS for now: MIPSVFPUUtils.cpp's
ReadVector/WriteVector/ReadMatrix/WriteMatrix/VFPURewritePrefix -
these are shared with every JIT backend's compile-time VFPU code, so
parameterizing them would balloon this into a JIT-wide refactor. This
is a partial refactor; that's the next boundary to push on.

Several JIT backends (x86 Jit.cpp, ARM/ArmJit.cpp, ARM64/Arm64Jit.cpp,
x86/X64IRJit.cpp, RiscV/RiscVJit.cpp, LoongArch64/LoongArch64Jit.cpp,
ARM64/Arm64IRJit.cpp) bake the raw interpreter function pointer
directly into JIT-generated machine code as their "fall back to the
interpreter for this one op" mechanism, with only a single argument
register set up for the call. Rather than hand-editing register
allocation across four architectures that can't be build-tested here,
added MIPSInterpretTrampoline(MIPSOpcode op) - a 1-arg wrapper around
MIPSInterpret(currentMIPS, op) - and pointed all 7 such call sites at
it instead, leaving that codegen untouched. Two other call sites
(JitLogMiss, JitBranchLog) were plain C++ calls and just got the
extra argument directly.

Verified (Windows x64): PPSSPPWindows/PPSSPPHeadless/UnitTest all
build clean, 49/49 unit tests pass. `test.py -g --graphics=software`:
interpreter 312/314 (cpu/fpu/fpu is the pre-existing, unrelated
interpreter-vs-JIT denormal difference; gpu/rendertarget/copy passes
standalone, so was cross-test state bleed in the batch run, not a
regression), default JIT 314/314, jit-ir 313/314 (gpu/vertices/morph
is an expected difference from the vertex decoder taking a different
mode with this core change, not a bug).

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_019SKhm9wKEQzRUsx9mTrrtQ
2026-08-12 14:02:19 +02:00
Henrik Rydgård ba4c3d93f3 Merge pull request #22087 from hrydgard/interpreter-stuff
Interpreter: Use a generated switch-tree dispatcher
2026-08-12 14:01:32 +02:00
Henrik RydgårdandClaude Sonnet 5 27c1aa85ef Register InterpreterDispatch.cpp/.h in the UWP project; update AGENTS.md
Core/MIPS/InterpreterDispatch.cpp landed without being added to
UWP/CoreUWP/CoreUWP.vcxproj / .vcxproj.filters, breaking the UWP build
(CMake and the main Core.vcxproj build both succeed silently, so this
only surfaces when someone actually builds the UWP project). Verified
the fix by building CoreUWP directly via UWP/PPSSPP_UWP.sln.

Also updates AGENTS.md's "new .cpp/.c file" checklist from five places
to seven (adding the two UWP project files), documents that the UWP
build can actually be build-tested here via MSBuild rather than just
eyeballed, and notes this exact miss so it's easier to catch next time.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_019SKhm9wKEQzRUsx9mTrrtQ
2026-08-12 13:06:52 +02:00
Henrik RydgårdandClaude Sonnet 5 ff864be3e2 Make ExecInstruction's unhandled case a plain -1 return, not an embedded fallback
The generated dispatch tree previously fell back to the old
MIPSGetInstruction()-based slow path (via a shared goto label) for
anything it didn't recognize - both genuinely invalid encodings and
the handful of real instructions with no interpreter implementation
(tge/tlt/teq/...). That baked policy ("what to do when unhandled")
into mechanically generated code, which is the wrong layer for it.

ExecInstruction() is now honestly partial: every unmatched case
returns -1, and callers are responsible for handling that. The
generated file no longer calls back into MIPSInterpret()/
MIPSGetInstructionCycleEstimate() at all, and no longer needs
MIPSTables.h.

RunUntilFast()'s -1 handling also skips re-walking MIPSGetInstruction()
entirely: since ExecInstruction() is generated from those exact same
tables, a -1 can only mean "no MIPSInstruction::interpret for this
op" - MIPSGetInstruction() would just rediscover the same thing.
Extracted that shared "log + disassemble + assert + skip" behavior
into HandleUnknownInstruction(), called directly instead.

Verified with `test.py -g --cpu=interpreter --graphics=software`:
still 313/314 (same pre-existing, unrelated cpu/fpu/fpu failure).

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_019SKhm9wKEQzRUsx9mTrrtQ
2026-08-12 12:31:18 +02:00
Henrik Rydgård cc90546301 Prep for plumbing the MIPS context pointer into the interpreter. 2026-08-12 11:25:06 +02:00
Henrik Rydgård c138bbf482 Headless on Windows: Set up CRT to output abort() to stderr instead of dialogs 2026-08-12 11:09:41 +02:00
Henrik Rydgård 67108e4bdb Merge pull request #22079 from hrydgard/fix-debugger-core-review
Core/HW and Core/Debugger review fixes by Claude
2026-08-12 10:55:52 +02:00
Henrik Rydgård a6e9a11662 Merge pull request #22082 from hrydgard/fix-ui-data-review
UI review and fixes by Claude
2026-08-12 10:55:19 +02:00
Henrik RydgårdandClaude Sonnet 5 d000d5366a Wire the generated ExecInstruction dispatcher into the interpreter's hot loop
Adds Core/MIPS/InterpreterDispatch.cpp, the checked-in output of
GenerateInterpreterDispatch() (see the previous commit), and hooks it
into RunUntilFast() in MIPSTables.cpp in place of the old
MIPSGetInstruction()-based table walk + indirect call through
instr->interpret. The checked-with-breakpoints/memchecks path
(RunUntilWithChecks) is untouched for now, since it inspects
MIPSInstruction flags directly and correctness there matters most.

Also fixes a real crash in headless.cpp found while testing this:
cmdLineOptions.gpuBackend.value() would throw when unset (e.g. with
--graphics=software), now uses value_or().

Verified with `test.py -g --cpu=interpreter --graphics=software`:
313/314 pass; the one failure (cpu/fpu/fpu) is a pre-existing
interpreter-vs-JIT denormal (flush-to-zero) difference, confirmed to
fail identically with the old table-walking dispatch, so unrelated to
this change.

Adds Tools/update-dispatcher.py to regenerate InterpreterDispatch.cpp
from a built PPSSPPHeadless binary whenever the MIPSTables.cpp tables
change.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_019SKhm9wKEQzRUsx9mTrrtQ
2026-08-12 10:52:47 +02:00
Henrik RydgårdandClaude Sonnet 5 65495c76c0 Add a codegen tool to generate a fast switch-tree interpreter dispatcher
MIPSTables.cpp has walked a tree of tables on every single interpreted
instruction since forever, with a standing TODO asking for exactly this:
"generate smart dispatcher functions from above tables instead of this
slow method." GenerateInterpreterDispatch() does that - it walks the
same tables MIPSGetInstruction() walks at runtime, but resolves the
walk into a nested switch tree once, at generation time, with each
leaf calling straight into the existing MIPSInt::Int_* handlers and
returning that instruction's fixed cycle count. Anything not covered
(invalid opcodes, and the handful of instructions with no interpreter
implemented at all, e.g. tge/tlt/teq) falls back to the existing
MIPSInterpret()/MIPSGetInstructionCycleEstimate() slow path, so the
result is total over all 32-bit inputs, same as the table-walking path.

Wired up via a new headless --generate-interpreter-dispatch flag,
which prints the generated Core/MIPS/InterpreterDispatch.cpp source to
stdout and exits.

Also widens CmdLine.cpp's --help column formatting, which silently
truncated any option name longer than 24 characters - the new option's
name was the first to hit it.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_019SKhm9wKEQzRUsx9mTrrtQ
2026-08-12 10:24:01 +02:00
Henrik RydgårdandClaude Sonnet 5 d718bda178 Rename MIPSInt/MIPSIntVFPU to Interpreter/InterpreterVFPU
Just a file rename (plus updating all build configs and includes) to
give the interpreter source files clearer names. No functional change.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_019SKhm9wKEQzRUsx9mTrrtQ
2026-08-12 10:04:40 +02:00
Henrik Rydgård 9315c0953a GamepadEmu: bounds-check touch pointer IDs before use
TouchInput::id was used directly to index the global primaryButton[]
array (MultiTouchButton::Touch) and to shift pointer bitmasks
(PSPDpad/PSPStick/PSPCustomStick/GestureGamepad::Touch), guarded only
by a debug-only assert in one of the five call sites - a no-op in
release builds. input.id isn't always a small sequential slot in
[0, TOUCH_MAX_POINTERS): SDL assigns SDL_FingerID values directly,
Android pointer IDs can go up to 31, and UWP's TouchMapper allocates
one more slot (11) than TOUCH_MAX_POINTERS (10) and can also return -1
when it runs out of slots - all reachable through ordinary multi-touch
use, no malicious input required.

Also apply bounds check to the PER_GAME gesture config ints
(iDoubleTapGesture/iSwipeUp/Down/Left/Right) before indexing
GestureKey::keyList[] with them.

Additionally, minor cleanup on Android and moves the TouchMapper helper
out from UWP to InputState.h.
2026-08-12 09:47:13 +02:00
Henrik Rydgård 1830945ca7 MpegDemux: fix heap-underflow rewind and add hard bounds to read8/skip
demux()'s "not enough data, rewind and try again next time" logic
unconditionally subtracted 4 (or 6) from m_index, assuming that many
bytes were consumed scanning for a start code. But the inner scan can
also exit via reaching the end of the buffer without finding a start
code at all, having consumed fewer bytes than that - when the buffer
holds under 4 bytes total, m_index goes negative, and the subsequent
memmove(m_buf, m_buf + m_index, size) then reads before the start of
m_buf. Clamp the rewind to 0.

read8()/skip() also had no bounds check against m_len (the actual
buffer allocation) at all - readPesHeader()'s header-length fields are
only cross-checked against the outer PES packet length, not against
how much data is actually available, so a crafted stream claiming a
long header could walk m_index past the buffer. Bound both against
m_len directly, at the lowest-level primitives so every caller is
covered.
2026-08-12 09:43:23 +02:00
Henrik Rydgård 90d8be9348 sceSasSetGrain/SasReverb: validate grain size and reverb preset index
sceSasSetGrain took no validation at all, unlike sceSasInit's grain
size check - a bad value could both throw on SasInstance::SetGrainSize's
allocation and, for a moderately large but successfully-allocated
value beyond PSP_SAS_MAX_GRAIN, read out of bounds of the fixed-size
mixTemp_ buffer during mixing. Apply the same bounds sceSasInit uses.

SasReverb::SetPreset() only checked the upper bound of `preset` before
indexing presets[], not the lower bound (-1 means "off"). The only
live HLE entry point (sceSasRevType) already clamps to [-1, 8], but
DoState() passes a savestate-deserialized value straight through with
no revalidation, so a corrupted/malicious savestate could index
presets[] negatively.
2026-08-12 09:43:23 +02:00
Henrik Rydgård bfbe44ad1c SimpleAudioDec: fix unsigned underflow and unvalidated stream-data size
FindNextMp3Sync() computed `sourcebuff.size() - 2` as the loop bound;
when size() is 0 or 1 this underflows to a huge size_t, turning the
scan into an out-of-bounds read. Reachable via sceMp3NotifyAddStreamData
followed by sceMp3Decode with as little as 1 pending byte.

AuNotifyAddStreamData() trusted the game-supplied `size` outright: a
negative value would make sourcebuff.resize() attempt a huge
allocation (via size_t underflow), an unbounded positive value grows
sourcebuff without limit, and the validated range didn't match the
actual read range (checked [AuBuf, AuBuf+size) while reading from
[AuBuf+offset, AuBuf+offset+size)). Validate size is positive and
capped to the buffer's declared capacity, and validate the range
actually read.
2026-08-12 09:43:23 +02:00
Henrik Rydgård 029066a977 DisassemblyManager: fix zero-size-symbol crash, OOB read, and other bugs
- DisassemblyFunction/DisassemblyData::getLineAddress() indexed
  lineAddresses[0] unconditionally; a zero-size symbol (reachable via
  the WebSocket debugger's hle.func.add/hle.data.add with an
  attacker-controlled size, a crafted ELF symtab entry with
  st_size==0, or the debugger UI's "set function size") leaves that
  vector empty, making findDisassemblyEntry's getLineAddress(0) call
  undefined behavior. Fall back to the symbol's own base address when
  out of range instead.
- DisassemblyData::createLines() detected an invalid address range and
  logged it, but fell through anyway into a loop reading through that
  whole range with the Unchecked memory accessors, which on
  non-masked builds do a raw pointer dereference with no bounds check
  at all. Added the missing return.
- DisassemblyLineInfo::ToString()'s snprintf calls all used
  sizeof(text) where text is a char* parameter (pointer size, not
  buffer size), silently truncating all output to a few characters
  instead of using the real bufSize parameter that was passed in but
  never used.
- analyze()'s misaligned-tail-data case stored the DisassemblyData
  entry under key alignedNext, but constructed it with the earlier
  (possibly much earlier) `address` as its own base address instead of
  alignedNext, misattributing those bytes to the wrong location.
2026-08-12 09:43:23 +02:00
Henrik Rydgård 9afab39bc8 MemBlockInfo: fix double-lock deadlock in FindWriteTagByFlag (previous commit)
FindWriteTagByFlag(flush=false) is only called from
FormatMemWriteTagAtNoFlush(), which is itself only called from within
FlushPendingMemInfo() - which already holds pendingReadMutex for its
entire body. The previous commit added an unconditional lock of that
same (non-recursive) mutex here, so any path that reaches a memory
write's tag formatting while a flush is in progress double-locks it
and hangs/crashes.

Repro: PPSSPPHeadless --graphics=software on
pspautotests/tests/gpu/clipping/homogeneous.prx reliably hit this.

Only take the lock when flush=true (i.e. when we're not already
guaranteed to be called from inside FlushPendingMemInfo's locked
section), using a defer_lock so the two call sites stay consistent.
Verified fixed: 4/4 clean runs of the repro above, plus the full
UnitTest suite (49/49) still passes.
2026-08-12 09:43:16 +02:00
Henrik Rydgård 527c50d32e MemBlockInfo: fix unsynchronized access to the slab maps from readers
The background flush thread calls FlushPendingMemInfo() at any time
while the emulator runs, holding pendingReadMutex for its whole body
while calling MemSlabMap::Mark() - which does new/delete and relinks
the intrusive Slab linked list via Split()/Merge().

FindMemInfo()/FindMemInfoByFlag()/FindWriteTagByFlag() only acquired
that lock indirectly and conditionally, inside FlushPendingMemInfo()
itself when the requested range happened to overlap pending data - the
actual .Find()/.FastFindWriteTag() traversal that followed ran
completely unsynchronized against the background thread's Mark() calls
on the same maps. This is a genuine use-after-free: a reader could
dereference a Slab* the flush thread just deleted, or race on the
shared lastFind_ pointer both sides read and write. Since a Slab's tag
is copied into the debugger's/WebSocket API's response, this could
also leak stale/freed heap bytes back to a caller. MemBlockInfoDoState
had the same gap around allocMap/suballocMap/writeMap/textureMap's
.DoState() calls.

Hold pendingReadMutex for the duration of these calls too, matching
the comment already on FlushPendingMemInfo ("This lock prevents us
from another thread reading while we're busy flushing") which wasn't
actually honored by the reader side.
2026-08-12 09:42:32 +02:00
Henrik Rydgård 8f5e984238 DriverManagerScreen: reject zip-slip in custom driver install
Both meta.json's "name" field and the zip entry names from a custom
GPU driver zip were joined onto GetDriverPath() verbatim - Path::operator/
does no ".." normalization, and ZipFileReader doesn't sanitize entry
names either. A crafted driver zip (these are commonly
downloaded/shared from third-party sites, e.g. Turnip/Mesa driver
packages) could use a ".." component to write files outside the
intended drivers/<name>/ directory. Reuse the existing
HasParentDirComponent() helper to reject such names.
2026-08-12 09:29:31 +02:00
Henrik Rydgård 4239c29928 BackgroundAudio: fix crashes and OOB reads parsing WAV/AT3 files
raw_bytes_per_frame (the 'fmt ' chunk's blockAlign field) is
unvalidated file data, and was used unchecked in three places:
- Divided into the 'data' chunk size to compute numFrames - a value
  of 0 divides by zero (crash).
- malloc()'d for raw_data was never null-checked before ReadData()
  wrote into it.
- Passed directly as the read length to the audio decoder on every
  frame, regardless of how much data is actually left in raw_data at
  the current offset - a bogus blockAlign larger than the real 'data'
  chunk size reads past the (padded) allocation into the decoder.
  Clamp it to what's actually available.

IsSimpleWAV() only checked raw_bytes_per_frame's upper bound, not that
it exactly matched one of the two cases Sample::Load() actually
handles (16-bit or 8-bit raw PCM) - a value in between passed the
check but matched neither of Load()'s conversion branches, leaving its
output buffer uninitialized and played back as heap garbage.

Reachable via a WAV/AT3 file parsed by BackgroundAudio.cpp - either
the menu background music preview (any EBOOT.PBP's SND0.AT3 track,
just from browsing the game list) or a user-configurable achievement
sound file.
2026-08-12 09:29:17 +02:00
Henrik Rydgård e35764d8e0 Store: reject path traversal in the store index's "file" field
entry.file comes from the remote store catalog (index.json) and is
joined onto DIRECTORY_GAME verbatim in OnLaunchClick() - Path::operator/
does plain string concatenation with no ".." normalization. On
platforms/builds where the index isn't fetched over HTTPS
(SYSPROP_SUPPORTS_HTTPS false), a network MITM or a compromised store
backend could use a crafted "file" value to make "Launch" boot an
arbitrary host file path instead of the selected store item. Reuse the
existing HasParentDirComponent() helper to reject such entries.
2026-08-12 09:27:47 +02:00
Henrik Rydgård 7f7fe3c712 CwCheatScreen: fix crash on malformed cheat database lines
GetLineNoNewline() computed `line + strlen(line) - 1` to strip a
trailing newline; fgets() doesn't stop at embedded NUL bytes, so a
line starting with one made strlen() return 0, pointing `end` one byte
before the buffer (a small OOB read, and conditionally an OOB write if
that byte happened to equal '\n').

ImportCheats() also called substr(4) on a "_C" cheat-name line without
checking its length first - a line that's just "_C"/"_C0"/"_C1" with
no name is shorter than 4 characters, and substr() throws
std::out_of_range, uncaught anywhere in this call chain, crashing the
app.

Both are reachable by importing a shared/downloaded cheat database
file (the normal way users add cheats), which could be corrupted or
maliciously crafted.
2026-08-12 09:27:47 +02:00
Henrik Rydgård b68e0c0896 RetroAchievementScreens: fix use-after-free on leaderboard screen close
rc_client_begin_fetch_leaderboard_entries[_around_user]()'s returned
async handle was discarded instead of stored in pendingAsyncCall_, so
the destructor's "abort the in-flight request" guard was always a
no-op. Closing (or switching modes on) a leaderboard screen before its
network fetch completes let the completion callback run later against
the already-destroyed screen (writing to its now-freed
pendingEntryList_/pendingAsyncCall_ members) - ordinary usage on a
slow connection, not just an edge case.
2026-08-12 09:27:47 +02:00