puppet-and-renderer-notes.md 110 KB

Puppet and renderer investigation notes

Puppet attachment/cropoffset positioning bugs

Multiple workshop items show attached puppet sub-parts (hair, halo/"nimb", hands, eyes) rendering detached from where they should be - floating away from the parent bone they're meant to sit on. Confirmed present in koshini (3621923790, bodyhairkochuru hair), Arona (2901582518, halo), and others. This is NOT the same bug as the (already fixed, verified) missing-body Z-clipping and mirrored-face Y-flip issues in CImage.cpp's puppet position baking - those are solid and shouldn't be touched.

What's been ruled out, with evidence:

  • cropoffset cannot be reverse-engineered from exported mesh/bone data. Tried three formulas (vertex bounding-box center, area-weighted triangle centroid, root bone bind-pose position), each auto-derived and applied via localTransform()'s existing model->cropOffset mechanism (see Model.h's cropOffset field and ModelParser.cpp:33). Verified against wallpapers that already ship an explicit, correct cropoffset (e.g. 2901582518's Arona_puppet.json has cropoffset: "-773.5 0") - none of the three formulas landed anywhere close. cropoffset is very likely author-authored metadata (set by hand in Valve's editor) that's simply absent from assets that don't declare it; there's no way to recover it from the packaged files alone. Don't retry this approach without new information.
  • Bone-track fallback isn't the cause for koshini's hair specifically. updatePuppetSkinning() has a documented fallback (see its comment around line ~1296) where a bone missing an animation-clip track falls back to its raw bindLocal position, which can be "thousands of units off" and visibly detach whatever it drives - a real, previously-observed failure mode. Checked via a temporary per-bone diagnostic (m_boneTrackDiagLogged in CImage.h/.cpp - still present, one-shot log per object, harmless to leave): all 39 of bodyhairkochuru's bones have valid matching tracks, and bone 0's animated position exactly equals its bind pose. Not what's happening here.
  • Not a missing-geometry / wrong-mesh-block-selection bug. resolvePuppetVertexLayout's coherence-scoring heuristic (which picks a vertex stride/candidate block by trying every plausible one) looked suspect since the bind-pose vertex position bounding box didn't match the texture's actual alpha content bounding box for bodyhairkochuru. Turned out to be a red herring: cross-checked the mesh's UV coordinate range against the texture's alpha bbox directly (tools/decode_tex.py + manual alpha thresholding) and they match closely. A puppet mesh's bind-pose vertex positions legitimately don't need to resemble its UV layout - those are two independent mappings (position = rest shape in mesh-space, UV = texture sampling), so this mismatch was expected, not a bug.
  • "Bind-pose vs. animated-pose confusion" (a prior session's leading theory for cropoffset) is superseded, not confirmed. The reasoning was sound (cropoffset formulas were computed from unskinned .mdl positions while the render path prefers skinned positions once animation is active), but it was never implemented, and the actual root cause found afterward (below) lives in a completely different code path (attachment resolution, not cropoffset). Don't rule it out for some other still-unexplained case, but it's no longer the leading hypothesis for the hair/halo drift.

Root cause found and fixed (2026-08-10): disassembled the real Wallpaper Engine renderer (wallpaper64.exe, via IDA Pro's idalib headless Python API + Hex-Rays) to see how it actually resolves an attachment point, rather than continuing to guess from this port's own code. Found and decompiled the exact native functions backing the scripting API's getAttachmentMatrix/getAttachmentOrigin/getAttachmentAngles/getAttachmentIndex (all registered together in one "Object" class-registration function, string-xref'd from "getAttachmentMatrix" etc. - sub_14016DC00 binds them to sub_14016D960/sub_14016DA30/sub_14016DB00/sub_14016D940 respectively). The real algorithm is a full 4x4 transform-matrix multiply: attachmentWorldMatrix = thisObject.getTransformMatrix() * boneLocalMatrix(attachmentIndex) (confirmed: getTransformMatrix's own native impl, sub_14016D0C0, is the exact same "this object's own matrix, no params" call the attachment getters make; the actual multiply is sub_1400BBD80, a generic 4x4x4x4 matmul with every rotation cross-term present, not a translation-only shortcut) - then origin/angles are the translation/decomposed-rotation of that combined matrix (sub_14016B6B0 does the atan2-based Euler decomposition for the angle getter).

This port's CImage::resolveTransform() (the chain[i]->attachment.has_value() branch) was only ever doing the translation part of that (anchorOrigin = parentOrigin + boneMeshPosition * parentScale, no rotation anywhere), and never fed the bone's own rotation into the attached child's angle at all - there was even a leftover one-shot TEMP-DIAG attachment rotation for head diagnostic already computing boneAngleDeg via the same atan2 trick, apparently once suspected but never wired anywhere. Fixed in CImage.h/.cpp: getAttachmentPointMeshPosition became getAttachmentPointMeshTransform, returning both position and the bone's current angle (std::atan2(animatedWorld[0][1], animatedWorld[0][0]), same mesh-space convention as PuppetAttachmentPoint::localTransform); resolveTransform now rotates the mesh-position offset by the parent's own accumulated resolved.angle before adding it (previously assumed the parent never rotates), and composes an anchorAngle = resolved.angle - meshTransform->angle (negated because mirroring the Y axis, which the existing code already does for the position, negates a rotation too) that both the attached child's own local-origin offset and its final accumulated angle now use instead of bare resolved.angle. When the parent has no rotation and the bone's local matrix has no rotation, this reduces to exactly the old formula (mathematically a strict generalization, not a rewrite) - so it shouldn't regress any case that was already correct.

Verification status, confirmed on the user's real machine (2026-08-10) after this session's fix landed:

  • 2986377055 - fixed. Confirmed by the user. (Puzzling in one respect: this wallpaper's parts rig as siblings under a shared group transform, not via the attachment mechanism this fix touches - per the sandboxed headless test's own logs, no Attachment resolve for ... line ever appeared for it. Either the sandboxed test missed an attachment path that's real on actual hardware, or the improvement here came from something else this fix incidentally also touched - e.g. the parent-rotation term added to the non-attachment offset math too. Worth understanding properly rather than assuming, next time someone's in this code.)
  • 3771392318 (spirit-blossom-ahri) - partially fixed. The orb attachment (ahriorb2, point orb) "looks better" per the user. The aaaaaaaaa attachment (ahriarm, her hand) is still wrong - looks 180°-flipped or mirrored - but the user confirmed this exact symptom was present identically before this fix too, so it's a separate, pre-existing bug this fix doesn't touch and didn't regress. A mirror/reflection is a different kind of transform than a rotation (not decomposable as a single angle), so it's very unlikely to be fixable by extending the same getAttachmentPointMeshTransform angle-only approach - if picked up again, look at whether ahriarm's own mesh/bone data (or the aaaaaaaaa attachment point's localTransform) has a negative-determinant (mirrored) component that nothing in the current pipeline (bind pose or this fix) accounts for.
  • 2901582518 (Arona) - still wrong, but the earlier "entire body renders as a white square" finding from this session's own headless/software-rendering (Xvfb + llvmpipe) test was WRONG - retracted. The user confirms Arona's body renders normally on real hardware; only some objects (the halo/nimb, presumably, matching the original report) float incorrectly. The white square was very likely a software-rendering (llvmpipe) artifact specific to whatever effect/shader Arona's scene uses, not a real engine bug - don't waste time chasing it again without independently re-confirming on hardware first. Her halo doesn't use the attachment mechanism at all (no Attachment resolve for ... line appears for this wallpaper in the sandboxed test either) - it very likely still goes through cropoffset, meaning the original, still-unimplemented "cropoffset needs skinned not bind-pose positions" theory from earlier in this doc may be the more relevant lead for Arona specifically, not this session's attachment-rotation fix.
  • 3621923790 (koshini) - unfixed, no visible change at all. Matches what the sandboxed test predicted (the hair's head/h attachment bone reported 0 rotation at every frame sampled) - this fix's rotation term is genuinely inert for koshini's specific rig, not merely unconfirmed. The root cause for koshini's hair drift remains open; needs a fresh angle (no pun intended) - re-examine whether it's positional (translation-only, unrelated to rotation) rather than rotational.

Real root cause found for a related-but-distinct bug class: cropoffset double-applied to non-puppet overlays, and missing paint-order (sortorder)

While chasing the attachment/animation-layer bugs above, the user reported a different wallpaper (3221531573, "asagi") with its eyes invisible and one hand rendering as a disembodied floating piece elsewhere on screen. This turned out to be two distinct, now-fixed bugs, unrelated to puppets/bones/attachment entirely - "asagi" has none of that, just plain autosized image layers with cropoffset (no puppet field at all in any of its models/*.json).

Bug 1 - cropoffset was being double-applied for non-puppet objects. localTransform()'s image->model->cropOffset handling (the same mechanism discussed above for puppets) was applied unconditionally to every autosized Image, puppet or not. Empirically confirmed (via disabling it entirely, rendering, and comparing pixel-for-pixel against the workshop item's own preview.gif - a very effective ground-truth check, worth doing more often) that a plain non-puppet autosized overlay's declared origin in scene.json is already the final correct placement - cropoffset must already be baked into it by the editor/exporter for these, and adding it a second time sends the layer flying off to roughly double the intended offset. With cropoffset disabled entirely, asagi's eyes and hand landed exactly where they should, pixel-matching the preview. But disabling it unconditionally breaks the puppet case: re-tested Arona (2901582518), whose Arona_puppet.json has an explicit cropoffset - with the addition removed globally, her entire puppet went off-screen (previously it was at least partially visible). So the fix is conditional, not a removal: localTransform() now only applies cropoffset when image->model->puppet.has_value() - i.e. only for objects that actually have a bone rig, matching the very first hypothesis for what cropoffset is for ("a puppet piece attached at the wrist, with the rest of its mesh extending away from center"). Verified: asagi's eyes/hand now pixel-match its preview.gif, and Arona's puppet bounding box/position numbers are unchanged/plausible (still can't visually confirm Arona due to the same sandboxed llvmpipe rendering limitation noted elsewhere in this doc, but the logged transform numbers before and after are consistent).

Bug 2 - paint order (sortorder) was never implemented at all, and "fixing" it with an id-based default caused real regressions - reverted to a safe no-op. Confirmed by grepping this whole codebase (zero matches outside vendored Catch2) that sortorder was never read. Disassembling wallpaper64.exe found a real, dedicated mechanism: a scene-level customsortorder flag and transparentsorting flag (registered right next to camera settings in sub_140153E40), plus a per-object sortorder int property (registered alongside the attachment getters in sub_14016DC00). asagi's scene.json sets neither flag and no object declares an explicit sortorder.

First attempt: sorted every object by sortOrder.value_or(id) (id ascending as an inferred, not disassembly-confirmed, default - the real per-frame sort/compare function was never located). This visually "fixed" asagi's eyes in isolation, but shipping it broke multiple other, previously-correct wallpapers on the user's real hardware: eyes disappearing again on 3528590419, Arona's entire puppet going invisible (screen just showing a faint blueish background), and 3764765600 rendering low-res. This makes sense in hindsight - these scenes have compositing passes (BG_redraw, Full Composition Layer, bloom, etc., visible in the object list) whose correctness likely depends on being processed in array order for reasons beyond simple front-to-back visual stacking (e.g. which pass reads which intermediate FBO) - reordering by id broke that, even though it happened to look right for asagi's specific case.

Reverted the default to each object's own array position instead of its id (sortOrder.value_or(arrayIndex)) - this is a mathematically exact no-op for every wallpaper that doesn't declare an explicit sortorder (the vast majority, apparently including all of the wallpapers touched this session), since sorting by array index reproduces array order exactly. Re-tested all four regressed wallpapers plus asagi after this revert: asagi's eyes/hand are still correctly shown (see below - it turns out this was never actually a paint-order problem for asagi, only a cropoffset one, see next paragraph), and the four regressions all look restored (koshini's hair back in front instead of behind; 3528590419's eye puppet objects drawing successfully per its log; 3764765600 sharp, not low-res). Object::sortOrder and the explicit-sortorder-respecting code path are still there (harmless, and correct per the disassembly for the rare wallpaper that does set it) - only the default when absent changed back.

A likely structural root cause found for the still-open attachment position bugs (koshini's hair, ahri's arm, and a newly-found case: mikasa's eye on 3764765600) - not yet fixed, high risk to fix carelessly. Investigated 3764765600 ("mikasa"), whose mikasa eye object attaches via attachment: "eye" to mikasaback's puppet rig - real ground truth (the item's own preview.gif) confirms the eye should be clearly visible near the face; in this port it lands near the very top edge of the screen instead. Ruled out several things first: not caused by the multi-layer animation blending work (confirmed via an isolated single-layer-only test rebuild - meshPosition/boneAngleDeg were bit-for-bit identical with 1 vs. 14 active layers), not a bone-track fallback (hasTrack=1 for all 73 of mikasaback's bones), not an alignment issue (neither object sets one).

Disassembled sub_140148A20 in wallpaper64.exe - the real engine's actual "compute this object's cached world transform matrix" function (found via the setParent-reconciliation function, sub_14016C6D0, which calls it repeatedly). It builds parentWorldMatrix * translate(origin.x, origin.y, origin.z) * scale/rotation, reading origin.y completely raw with no sign flip anywhere in that composition - the real engine evidently uses one consistent Y convention all the way from JSON through every level of parent/child matrix composition, with any Y-flip happening at most once, at the very end (camera/projection), not interleaved between objects.

This port's pipeline instead applies Y-flips at multiple separate points: updatePuppetPositionBuffer's local-canvas bake (size.y/2 - source[index+1]), the CImage constructor's/updateScenePosition's scene-height flip (sceneHeight/2 - m_pos.y), and the attachment formula's own -meshTransform->position.y negation (this session, mirroring what the pre-existing code already did for the non-rotated case). Each of these is self-consistent for a single object's own rendering (which is presumably why simple single-object cases and small attachment offsets mostly look right), but the attachment path bridges two objects through this multi-flip pipeline - if the number of flips a value passes through differs between the "parent's mesh-space bone position" side and the "child's own screen-placement" side, the result is a real, compounding Y error that gets worse the larger the offset - matching what's observed: asagi/small attachments looked closer to right, ahri's orb "looked better" after this session's rotation fix, but mikasa's eye (a large offset with a large bone rotation) is badly wrong on the Y axis specifically, while X looks roughly plausible.

Fix attempted and looking strong in the sandbox (2026-08-10), pending real-hardware confirmation. Rather than touch the whole pipeline, made the smallest change consistent with the sub_140148A20 finding: in resolveTransform's attachment branch, stopped negating meshTransform->position.y and stopped negating meshTransform->angle - i.e. the bone's mesh-space position/angle now folds into the (already-unflipped, confirmed by how ordinary non-attachment children fold in two lines below with no negation) origin-space chain raw, the same way a normal child's local.origin does, instead of getting an extra Y-flip that only makes sense for the different mesh-to-screen baking pipeline (updatePuppetPositionBuffer) it was borrowed from.

Result on all four attachment-position bugs open at the start of this fix:

  • koshini (3621923790) - looks fixed. Hair now renders correctly on the head, no more floating/detached strand, in the sandboxed render.
  • mikasa (3764765600) - looks fixed. The eye attachment moved from anchorOrigin=(2050.62, 83.26) (top edge, previously invisible) to (2050.62, 1356.74), and now renders clearly on her face, matching preview.gif.
  • Arona (2901582518) - unaffected, as expected. Her root puppet has no attachment (no parent), so this branch never runs for it; logged transform is bit-identical before/after. Her halo bug is still whatever it was (not this).
  • ahri (3771392318) - position changed (ahriarm's anchor moved by roughly (-1450, +780)), but couldn't visually confirm the arm's mirror/orientation is now correct - spiritblossomahribase renders as a flat black silhouette in this sandbox (the same LightingV1-shader llvmpipe limitation noted elsewhere in this doc), so the arm itself isn't visible enough to judge by eye here. Needs a real-hardware check.

Not yet confirmed on real hardware - that's the immediate next step. This reverses a sign convention two things depended on (a previous session's translation-only formula, and this session's own earlier rotation-composition work built on top of it), so it's a meaningful change - but it's also a single, well-reasoned, disassembly-grounded fix (not a broad rewrite), reduces to a plausible no-op for cases where meshY≈0, and produced two clear, visually-obvious wins in the sandboxed test with no observed regression. If it turns out wrong on hardware, the negated version is in git history right before this change.

Partial correction: paint order was never needed for eye visibility, but does affect a real, minor compositing seam. Re-tested asagi with the paint-order default fully reverted to array order (eyes-before-background) - the eyes still rendered correctly, on top, matching the preview.gif for basic visibility/position. But the user spotted a follow-up artifact: two faint ring/seam outlines around each eye that weren't there when paint order was id-based. Traced this precisely: decoded asagi 2's (the full-background art) alpha channel at the eye location and confirmed it has genuinely transparent, soft-edged (feathered, not hard-cut) eye-shaped holes for the overlay to show through. Empirically confirmed (temporarily forcing id-order back on, screenshotting, reverting) that the ring only appears with array order and disappears completely with id order - so paint order does matter here, just for compositing/blend-seam quality at a soft alpha edge, not for whether the eyes show up at all.

This leaves a real tension: id-order removes asagi's seam but severely broke several other wallpapers (invisible puppets, low-res output) when it was the default; array-order fixes those but leaves asagi with a minor cosmetic ring. Kept array-order as the default - a thin seam on one wallpaper is a far better tradeoff than broken rendering on several others. The real engine's preview.gif shows no seam at all, meaning WE's true default sort algorithm isn't simply "by array index" or "by id" - it's something else this session didn't identify (a transparentsorting-aware depth sort is a plausible guess, given that flag exists specifically for this kind of transparency-ordering scenario, but this is speculation, not confirmed). Whoever revisits this: the asagi ring is a small, safe, reproducible repro case for testing a better default (temporarily hardcode object->id as the sort key in CScene's constructor, rebuild, and diff the eye-area alpha seam against array order, same as this session did) - but don't ship a new global default without testing it against all four of the wallpapers that regressed this session, not just asagi.

Worth remembering generally: verify a theory in isolation (e.g. test the fix for JUST one hypothesized cause) before shipping a broader, riskier change built on top of it - the original id-order change bundled "fix visibility" (unnecessary, as it turned out) with "fix the seam" (real but narrower), and shipping both together is what caused the regressions.

Third default tried and also reverted: sort by object footprint (area). After id-order and array-order both proved insufficient defaults (one over-fixes and under-fixes depending on the wallpaper), tried a different, principled-looking heuristic: larger Image objects paint first/behind, smaller ones paint later/on top, tiebreaking on array position for objects with no static size (particles, FBO-autosized compositing passes, etc). Motivation was real: in at least one wallpaper's own data, object sizes strictly decrease with intended layering depth (body > face > eyes > pupils > highlights), and Arona's BG_redraw (the object whose very high id broke everything under id-order) is also, unsurprisingly, exactly full-scene-sized - both cases looked like they'd fall out naturally from a size-based rule.

It fixed asagi's seam (confirmed clean). But it broke something new: koshini's (3621923790) hair puppet piece disappeared entirely - confirmed via a direct pixel crop comparing before/after, not the pre-existing unrelated broken-texture issue that wallpaper also has (the user specifically flagged that distinction and asked to double-check, which is what caught this - a good instinct: don't assume a rendering artifact fits a theory just because it could, verify it against a known baseline). Reverted immediately back to plain array order.

At this point, three attempted defaults - id, array, and size - have each fixed at least one case and broken at least one other (except plain array order, which hasn't been observed to break anything, only to leave asagi's minor seam unfixed).

Fourth attempt: attachment/parent chain depth - safe, real improvement, but doesn't fix asagi's seam (her objects have no parent relationship to derive anything from). Diagnosed why size-order broke koshini specifically first, rather than abandoning size as a direction outright: her hair puppet piece (bodyhairkochuru) has a declared canvas of 3241x1932 (6.26M px^2) - larger than her entire body (2222x2222, 4.94M px^2), because a puppet's JSON size reflects its mesh's own local bounding box (which can sprawl for flowing hair/animated parts), not visual screen prominence. Sorting by that size put the hair behind the head that's supposed to sit in front of it.

But koshini's actual structure - koshinibody (root) -> koshirohead (attached to body) -> bodyhairkochuru (attached to head) - already encodes the correct paint order directly: parent-before-child, deeper-attached-object-last/on-top. This is real structural data (parent/attachment), not a proxy like size or id. Implemented: primary sort key is attachment-chain depth (root = 0, one level of parent = 1, etc., capped at 32 like resolveTransform's own cycle guard), falling back to array position as the tiebreak for equal-depth objects - which is the common case (most wallpapers' objects are all at depth 0, so this is a no-op for them, same safety property plain array order has).

Tested against every wallpaper that mattered this session: koshini's hair intact (confirmed via pixel crop), mikasa's eye still visible, 3528590419's eye puppets still draw (glError=0), Arona's Transform for numbers bit-identical to the working baseline (both her key objects are depth-0 - this reduces to exactly plain array order for her), ahri's attachment numbers unchanged. No regressions found. Does NOT fix asagi's seam - asagi 1 (eyes) and asagi 2 (background) are both top-level with no attachment/parent relationship to each other, so depth-sort can't differentiate them and falls back to the same array order that already has the seam. Real, safe, additive value on its own (any other wallpaper with puppet attachment chains whose scene.json array order doesn't happen to match visual layering would previously have rendered wrong and had no fix - now it's structurally guaranteed correct regardless of array order), independent of whether asagi's specific issue ever gets solved. Not yet confirmed on real hardware.

Recommend NOT attempting a fifth guess at a different default for the asagi-seam case specifically without a fundamentally different source of ground truth - id, array, size, and depth have all been tried now, each grounded in real reasoning, and none solve asagi without breaking something else except depth, which is safe but structurally can't apply to her case (no parent relationship exists to use). If this remains worth chasing, the next idea should target that specific gap (two top-level siblings with an implicit "one is a background hole cut for the other" relationship) rather than another global sort-key guess.

Whether this also explains Arona's still-open halo bug is unknown - a scene object literally named "halo"/"nimb" wasn't found in 2901582518's scene.json at all (only bundled, possibly-unused materials/particle/halo_*.json particle-effect assets were found via package listing); the halo may be a particle system (a completely different rendering path, CParticle not CImage, that neither this fix nor the attachment-rotation fix above touches) rather than an image layer. Worth checking what object type actually renders Arona's halo before assuming either fix applies to it.

Net takeaway: the attachment-rotation fix is real and helped in at least one confirmed case, but it was not (and was never expected to be, per the disassembly - it only adds a rotation term) a fix for every reported instance of this class of bug. At least two more distinct bugs remain: a mirror/reflection issue (ahri's arm) and whatever's actually wrong with koshini's hair and Arona's halo (still unknown - possibly the old cropoffset/skinned-pose theory, possibly something else entirely). Don't assume they're the same bug just because the symptoms all get described as "floating"/"detached."

Follow-up (same session): scale/mirror propagation added. After the rotation fix landed, the user reported ahri's aaaaaaaaa-attached hand moved to the correct position (confirming the rotation fix's dominant effect was real) but still renders "180° flipped or maybe mirrored," identically to before the fix - i.e. a separate, pre-existing bug, not a regression. This matches exactly what the fix's own scope predicted it wouldn't cover: getAttachmentPointMeshTransform only ever decomposed the bone's matrix into position + rotation angle, never scale - and a reflection (mirrored bone, e.g. a rig that reuses one arm's mesh flipped for the other) isn't representable as a rotation angle at all. Extended AttachmentPointTransform with a scale (glm::vec2) field: scale.x is the transformed X-basis vector's length; scale.y is recovered as det(X,Y)/scale.x (projecting the transformed Y-basis onto the Y-direction a pure, unreflected rotation by the already-computed angle would have produced) - this comes out negative exactly when the bone's matrix includes a reflection, instead of silently folding that reflection into a bogus rotation angle the way a naive atan2-only decomposition (the first version of this fix) does. resolveTransform now multiplies this into the attached child's local.scale before folding it into resolved.scale, consistent with the "scale then rotate" composition order already used everywhere else in this function (see the offset-rotation line right below it). Builds clean; sandboxed test again shows boneScale=(1,1) at the one-shot diagnostic's sample frame for ahri's arm (same limitation as before - that log only ever captures the first attach, and this rig's mirroring, if that's really what's happening, may only be static/baked into the bind pose or only visible partway through the clip) - not yet confirmed on real hardware, and the user's follow-up report was "nope, ahri did not get better" - so the scale/mirror fix above did not fix ahriarm's orientation. Kept in the tree (it's still a real, disassembly-motivated generalization and hasn't shown any regression), but it is NOT the fix for this specific symptom.

Further IDA investigation (same session, following user's request to dig deeper): Traced the real engine's setParent()-with-world-preservation function (backs the setParent scripting method) and found its scale decomposition is unconditionally sqrt(sum of squares) per basis vector - no determinant/sign check anywhere. This is real evidence the real engine may not represent a mirrored bone as a signed scale component at all (undermining the scale/mirror fix's premise, though that function may not be the relevant code path for bone/attachment rendering specifically - it's the scripting API's reparent-while-preserving-world-transform operation, not necessarily the same decomposition used for puppet bones). Also searched for the runtime multi-animation-layer blend function extensively (checked the largest function in the binary that calls the matrix-multiply helper, 48KB, no clear per-layer-weighted-sum pattern found) - did not locate it. This binary has essentially no symbol names, making an undirected search increasingly unproductive; a future attempt should have a better anchor (e.g. a RenderDoc-style dynamic trace isn't possible since the original is a Windows binary that can't run in this Linux sandbox, but breakpointing the equivalent code path under Wine on a real Windows/Wine setup, if ever available, would be far more direct than static analysis alone).

Empirical experiment tried instead (per user's choice): multi-layer additive animation blending, now implemented. ahriarm's scene.json declares two simultaneous animationlayers ("Animation 1", "Animation 2", both additive: true, blend: 1.0) - and CImage's loadPuppetMesh/updatePuppetSkinning previously only ever played the first matching layer (a deliberate prior-session tradeoff, since naively summing every active layer's raw rotation over-rotated badly on rigs with many layers). This session:

  • Removed the early break in loadPuppetMesh so every matching, layer gets added to m_puppetActiveAnimations (already a vector, this was the only thing forcing it to one entry).
  • Rewrote updatePuppetSkinning to sample every currently-visible active layer independently (each can have its own rate/duration/frame), then per-bone accumulate a blend-weighted delta from a shared baseline (bind position, zero rotation, unit scale) using each layer's own ImageAnimationLayer::blend value (0-1) - previously read from the model but never actually used anywhere. For the single-active-layer, blend=1.0 case (every puppet this codebase already had working) this reduces to exactly the old formula, so it's a strict generalization, not a rewrite.
  • Builds clean; sandboxed test confirms both layers are now active (activeLayers=2 in the bone diagnostic) and produces different, non-exploding bone rotations (roughly half the magnitude of the single-layer values, not doubled/compounded) on koshini and ahri alike, with no crash or degenerate geometry.
  • Not yet confirmed on real hardware. This is a broader, riskier change than the earlier two fixes - it touches every puppet with multiple animation layers, not just ones using attachment, so real-world testing needs to cover more than just these 4 wallpapers (watch for regressions on any puppet that has multiple animationlayers and previously looked correct with only the first one playing).

Sandboxed headless rendering (tools/headless_render.sh, Xvfb + llvmpipe) is not trustworthy for judging puppet-with-effects wallpapers. Beyond Arona's white square (see above, confirmed a software-rendering artifact - her body is fine on real hardware), ahri's spiritblossomahribase puppet renders as a flat black silhouette under this same headless setup, correlating with the many Resolving require module: LightingV1 log lines for its shaders - looks like whatever lighting pass those puppets' materials need doesn't produce correct output under llvmpipe. Treat this tool as good for confirming "does it crash / are the numbers in the diagnostics sane" but not as a stand-in for a real GPU when judging whether a puppet-positioning fix actually looks right - that needs the user's real hardware, same as this whole investigation's fixes have needed so far.

Reproduction/verification workflow used throughout: RenderDoc capture via tools/rdc_trigger.so + tools/rdc_dump.py (see above), comparing the object's m_pos/resolved-origin diagnostic logs (TEMP-DIAG final resolved origin, Attachment resolve for ..., Transform for ... - all still present in CImage.cpp, left in per explicit request) against actual rendered pixels. Manual coordinate-space arithmetic (converting between the engine's pre-flip Y-down "origin" convention and post-flip Y-up "m_pos"/screen convention) proved extremely error-prone across this investigation - prefer pulling ground truth via RenderDoc/screenshots over re-deriving the transform chain by hand.

Attachment rotation fix (2026-08-10): mikasa's eye, follow-up to the Y-convention position fix

After the Y-convention position fix above landed and fixed mikasa's (3764765600) eye position (moved from off-screen to correctly on her face), the user reported the eye's rotation still looked wrong - an oddly-angled sliver instead of a natural closed-eye contour.

Traced it: mikasa eye's own JSON declares angles: "0 -0 0.77786" (~44.6 degrees, an artist fine-tune), and the eye attachment bone's own rotation is ~-45 degrees. With anchorAngle = resolved.angle + meshTransform->angle (the same "no sign flip" rule the position fix established, applied to rotation too), these nearly cancel to ~0 degrees net rotation on screen - explains the reported symptom exactly.

First attempt (reverted): flipping the sign. Changed to anchorAngle = resolved.angle - meshTransform->angle so rotation got the opposite sign from position. This visually improved the eye's shape in an isolated crop, but shipping it broke the eye's position badly - the user reported "mikasa's eyes now completely disappeared". Root cause: anchorAngle wasn't only the final stored angle, it was also the angle used to rotate the child's own local-origin offset into the bone's frame a few lines below (rotateVec2({local.origin * resolved.scale}, anchorAngle)). Mikasa eye's own declared local origin is a large offset (~355 units in magnitude), so swinging anchorAngle from ~-0.5 degrees to ~+90 degrees swung that offset by nearly 90 degrees too, pushing resolvedOrigin from (1964.14, 1014.53) to (2392.67, 1269.17) - past the right edge of a 2560px-wide scene. Confirmed via direct log comparison (grep -i "Transform for mikasa eye" on both runs) and reverted back to addition.

Second attempt (also reverted): decoupling the final angle from the offset-rotation angle. CImage.cpp's resolveTransform briefly tracked two separate variables in the attachment branch: anchorAngle (bone rotation folded in, unchanged, still drives the position offset-rotation) and a finalAngle (the parent chain's own accumulated angle only, with no meshTransform->angle contribution). This looked like a natural eyelash contour in a sandbox crop and was shipped on that basis - wrongly. The user's actual real-hardware screenshot (/workspace/2.jpg) showed bare skin, no eye at all - not the sandbox's suggested improvement. Root cause of the mistake: mikasa eye.tex isn't a full eye graphic, it's a tiny set of accent marks (a highlight dot, a sliver of iris shading, a couple of lash strokes - confirmed via decode_tex.py, only ~0.7% of the 444x444 canvas is non-transparent) meant to overlay a specific closed-eye crease that's baked directly into mikasaback's own base texture (confirmed by decoding mikasaback.tex and cropping its face region - a plain closed-eye line, no color). Rotating the mesh swings those small, precisely-placed marks to different screen pixels even though the object's own bounding-box center doesn't move - finalAngle apparently rotated them off that tiny target entirely. The workshop item's own preview.gif (48 frames, checked at several points in the cycle) confirms the eye is meant to render as a clear, consistently-visible reddish-brown almond shape with a highlight - not something that should ever fully disappear.

Reverted back to the original single-anchorAngle formula (resolved = { local.origin, local.scale * resolved.scale, local.angle + anchorAngle }) - the last state the user actually confirmed was visible on real hardware, even though wrongly rotated (near-0 degrees net rotation from the ~44.6 degree local angle and ~-45 degree bone angle nearly cancelling, same root cause as originally diagnosed). Verified via log: resolvedOrigin=(1964.25,1014.66) (matches the position-fix baseline) and resolvedAngleDeg=0.0618 (matches the original "thin sliver" symptom, not off-screen, not a different wrong angle).

Status: rotation is unfixed, but visibility is restored to the last known-good state. Two attempts at a different rotation value each broke something further (position, then visibility) - a real fix needs to either explain why some different angle would still land the same tiny accent marks on the same crease (not just look plausible in one crop), or find that the crease/accent-mark relationship depends on something not yet accounted for (e.g. the bind pose vs. current animated bone pose diverging in a way that isn't just a sign/inclusion question). Do not re-attempt a rotation-only change without a way to verify it beyond a single sandboxed screenshot - both prior attempts passed that bar and both failed on real hardware.

Asagi's eye-ring seam: exact mechanism found, no safe fix yet (2026-08-10)

Continued the paint-order investigation (see the four-sort-key section above) with a different question - not "what's the right global default" but "what is this specific seam, mechanically". Unpacked 3221531573's scene.pkg and decoded asagi 1.tex (the eye detail overlay, 511x344) and asagi 2.tex (the full-canvas background, 2560x1440) with tools/decode_tex.py to inspect their actual alpha channels directly, rather than reasoning about paint order in the abstract.

Findings:

  • asagi 2 (background) has a genuine soft-edged, feathered alpha cutout exactly where each eye goes - confirmed visually (bg_alpha_crop.png: a soft black eye-shaped hole on a white/opaque field) and numerically (alpha values form a gradient across the hole boundary, not a hard 0/255 edge).
  • asagi 1 (eye detail) has a binary alpha channel - only values 0 and 255 appear anywhere in the whole texture (checked via np.unique). It's a hard-edged opaque rectangle (with transparent padding around it, not a soft eye-shaped cutout at all) that fully and generously contains the background's hole (opaque bbox 668-1078 x 306-549 vs. the hole's bbox 728-1033 x 426-502 - 60-120px of margin on every side).
  • Array order in scene.json lists asagi 1 (id 21) at index 0 and asagi 2 (id 16) at index 1 - i.e. the eye detail paints first/underneath, the background-with-hole paints second/on top. That's backwards from the usual "base layer behind, detail on top" pattern.

Because asagi 1 is a hard-edged opaque rectangle, when it's UNDER asagi 2, the background's soft hole is the only thing determining what shows: fully transparent center -> asagi 1 shows through cleanly (correct), fully opaque field -> background skin shows (correct), but the gradient annulus at the hole's edge alpha-blends the background's own edge-region color over the already-opaque asagi 1 pixel beneath it - producing a visible ring of background color superimposed right at the hole's feather zone. That ring is exactly the reported "2 circles around her eyes".

This also explains a detail noted earlier without a mechanism: per the memory notes, forcing id-ascending order (which happens to put asagi 2, id 16, before asagi 1, id 21) made the ring vanish - with the background drawn first/underneath and the hard-edged opaque eye rectangle drawn second/on top, there's no gradient blend left to produce a seam; asagi 1 just cleanly overwrites everything in its footprint, hole or no hole.

Why id-order still isn't a safe general fix, confirmed against a second wallpaper. Checked 3528590419 (one of the four wallpapers id-order broke) directly: its face/shadow/eye/pupil/highlight objects (脸, 阴影, 左瞳, 左高光, 左眼, ...) are all "blending": "translucent" too - so "sort translucent-blended siblings by id" (a narrower version of the id-order idea, scoped to blend mode) doesn't rescue this case either. Its array order already encodes correct depth (face -> shadow -> pupil -> highlight -> eye, each progressively more "on top", matching intent); its id values don't correlate with depth at all (face is id 344 despite being the backmost layer) - id-sorting scrambles it regardless of blend mode. So the earlier finding stands: neither id nor array index is a globally correct depth proxy across these wallpapers' own scene.json data - it looks like an artifact of how each wallpaper's original editor session happened to create/list objects, not a recoverable convention.

Implemented and shipped: a geometric per-pair heuristic, not a global sort-key change. CScene.cpp's constructor now runs a refinement pass after the depth-based sort above: for every pair of top-level (no parent/attachment), plain (non-puppet), non-utility (models/util/* excluded), fully translucent-blended objects with no explicit sortorder, if one's bounding box (origin +/- size/2) is fully contained inside another's and the container is within 10% of the full scene size (i.e. actually looks like a background, not just a coincidentally large sprite), the contained object is moved to draw immediately after its container. This is deliberately narrower than bounding-box containment alone: the "near full-scene container" requirement is what keeps it from matching a lens-flare glow sprite (mikasa's flare chain, much smaller than the scene) or 3528590419's "Audio bar" (a models/util/composelayer.json reactive layer, already excluded by the utility-path check) or mikasaback/aotback (excluded because one side is a puppet, whose declared size reflects mesh bounding box, not visual footprint - the same lesson learned from the earlier size-based sort failure).

Verified via Python simulation against every wallpaper touched this session before writing any C++ (asagi, mikasa, koshini, Arona, ahri, 2986377055, 3528590419) that this rule only ever fires for asagi's eye/background pair - every other qualifying pair in every other wallpaper is either already correctly ordered or excluded by one of the guards above. Confirmed by the user on real hardware: ring is gone, "looks awesome", no regressions found across their normal wallpaper rotation.

A wallpaper (1746890525, a single giant multi-effect puppet - fire/pulse/waterwaves/reflection/shake/shine/depthparallax, 13 passes) was briefly reported as "half looks like lava" during testing, but turned out non-reproducible - the user was hotswapping heavily at the time, which lines up with the use-after-free below rather than this paint-order change (this wallpaper has only one qualifying Image, and it's excluded for being a puppet, so the containment rule is a structural no-op for it either way). Before it was confirmed transient, g_TextureNResolution uniform binding was checked directly via a one-shot diagnostic log in CPass.cpp (TEMP-DIAG texture resolution uniform for object 13, left in place) and found correct (1:1 ratios everywhere) - ruling out a UV/padding bug in the fire effect's flow-map sampling as the cause, in case this resurfaces.

Hotswap use-after-free crash (2026-08-10)

Real crash (SIGSEGV + coredump) reported on real hardware: CSound::applyEffectiveVolume() <- CScene::setAudioPolicy() <- WallpaperApplication::applyAudioPolicy() <- checkHotswapRequest(), alongside a filesystem_error: Cannot find requested file ... [/models/bar.json] for a different, apparently-incomplete workshop item (2352180009) hotswapped around the same time.

Root cause: both checkHotswapRequest() and advancePlaylist() moved the outgoing wallpaper's Project into a try-block-local variable, on the assumption it only needed to outlive the setWallpaper() call a few lines later (the new CWallpaper holds a raw reference into its Project's data, not a copy). But if anything between that move and setWallpaper() throws - e.g. CWallpaper::fromWallpaper() failing to load an asset the new scene references, exactly what a missing models/bar.json would do - the catch block is reached with setWallpaper() never having run: the old CWallpaper for that screen is still installed and still rendering every frame off data the local variable is about to free at scope exit. The next thing touching every installed wallpaper was applyAudioPolicy() (called unconditionally after the per-screen loop in checkHotswapRequest(), and inline in advancePlaylist()), which dereferenced the now-dangling Sound reference inside CSound::applyEffectiveVolume() and crashed - matching the reported stack trace exactly.

Torn/interlaced texture on koshini's hair bow decoration (2026-08-10, not yet fixed)

User-reported "buggy textures" on two wallpapers, both illustrated with real screenshots (/workspace/1.jpg, /workspace/2.jpg): koshini (3621923790) shows a horizontal-scanline "torn" look specifically on one wing of the butterfly-bow hair ornament near the crown; ahri (3771392318) shows a jagged black/orange/blue "hole" right at the ahriarm attachment seam (very likely the same still-open mirror/flip bug documented above, not investigated further this session - see that section).

Ruled out, with hard evidence, in this order:

  • Not a broken source asset. Decoded materials/bodyhairkochuru.tex directly out of the .pkg (tools/decode_tex.py, all 3 mip levels) - the drawn butterfly bow art is clean, no tearing, in the raw file. Whatever's wrong is introduced by this port's rendering, not present in what Valve shipped.
  • Not any of the five separate butterfly/bfly1/bfly2 scene objects. --list-objects reveals koshini's scene has decorative extras beyond the literal butterfly (id 249) object found by name search: butterfwly (148, hidden), bfly1 (304, 322), bfly2 (350, 336, both with "scale": "-1.08142 1 1" - a genuine negative-X mirror, initially suspected as the cause). Disabled every one of these (--disable-object) individually and all together - the torn wing was pixel-identical in every case. The bow is baked directly into the bodyhairkochuru puppet's own texture/mesh, not a standalone sprite.
  • Confirmed it IS bodyhairkochuru (id 134). --disable-object 134 removes the torn wing along with the rest of the hair, proving this object (not something layered under/over it) draws it.
  • Not anisotropic filtering. Temporarily forced GL_TEXTURE_MAX_ANISOTROPY to 1.0 in CTexture.cpp - artifact unchanged.
  • Not mipmap/LOD selection at all. Temporarily forced GL_TEXTURE_MIN_FILTER to plain GL_LINEAR (no mipmapping whatsoever) - artifact unchanged. Rules out the classic "mirrored-UV-seam mip popping" theory entirely.
  • Not any effect shader. koshini's hair has three effects (glitter, blur x4-pass gaussian, waterwaves). --render-debug skip-effect=324 (blur alone) and --render-debug base-only (every effect off, base material pass only) both leave the artifact untouched - it's present in the raw base-material draw, before any effect FBO pass runs.
  • Not a vertex-layout/stride misparse. Suspected resolvePuppetVertexLayout's coherence heuristic (see the attachment-bug section above - it only scores triangle position coherence, never validates UVs) might have picked a stride that reads garbage UVs for this one sub-region while still looking fine position-wise. Added a temporary diagnostic (TEMP-DIAG vertex layout candidates for bodyhairkochuru in CImage.cpp, still in the tree) dumping every candidate (block, stride) and its score: the winning stride=80 candidate scores 0.0198, more than 2x better than the runner-up (0.045) - not a close call. Manually parsed the raw .mdl bytes at that stride for several vertices by hand: blend weights sum to exactly 1.0, UVs all land in a plausible [0,1]-ish range - the vertex format identification is correct.
  • Not a per-triangle UV distortion either. Added a second temporary diagnostic (TEMP-DIAG uv/pos area ratio for bodyhairkochuru, still in the tree) computing every triangle's UV-space-area-to-position-space-area ratio and flagging any that deviates >20x from the mesh's median. Zero outliers across all 1592 triangles - no single triangle has a locally-wrong UV mapping.

Real lead found, not yet chased to a fix. Wrote a one-off script parsing the raw .mdl bind-pose vertex block directly (stride 80, matches the engine's own resolved layout) and bucketed all 959 vertices by rounded XY position to find near-duplicates. Found genuine pairs of vertices sitting within ~1-2 units of each other in bind-pose space but with substantially different UVs (e.g. vertex 34 at (-293.4, 474.4) uv (0.717, 0.144) vs. vertex 772 at (-292.8, 473.6) uv (0.410, 0.255), only 1.02 units apart). This is consistent with how flowing-hair meshes are normally authored - many separate strand quads that legitimately converge near the scalp in bind pose before spreading apart under animation, each sampling a different narrow strand region of the texture - so it isn't obviously a bug by itself. But bodyhairkochuru.json's material declares "depthtest": "disabled" and every vertex's Z is a flat 0.0, so nothing separates overlapping layers by depth; if the animated (skinned, not bind-pose) screen positions of two such originally-separated strand quads happen to converge at the wing's location at some point in this rig's animation loop, they'd alpha-blend on top of each other in whatever order they're drawn, which is a plausible mechanism for a scanline-like "torn/double-exposure" look on a high-contrast decal (the bow) while staying invisible on the flat-colored hair strands elsewhere. Not yet confirmed - this would need comparing animated/skinned screen-space positions (not the static bind pose checked so far) across the rig's actual bone weights for the specific vertices that render as the bow, which wasn't done this session.

Two temporary one-shot diagnostics were left in CImage.cpp for whoever picks this up next: the vertex-layout-candidate dump and the per-triangle UV/position-ratio dump, both scoped to this->getImage().name == "bodyhairkochuru" so they're inert for every other wallpaper.

Reproduction is fast and doesn't need real hardware: tools/headless_render.sh against workshop item 3621923790 (needs --assets-dir pointing at Wallpaper Engine's real assets/ folder) reproduces the exact same torn pattern seen in the user's real screenshot, in the hair-bow region roughly at screen fraction (0.69-0.86, 0.02-0.23) of a 1920x1080 capture.

Follow-up, same day: found the actual overlapping geometry, and it's a second, unrelated hair-wisp element - but couldn't determine whether the overlap itself is a bug. Confirmed the "torn" look is present, in a very similar shape, at every sampled point across the animation loop (frames 5/15/45/90/150 of 180 all show it) - it isn't a transient animation-phase artifact.

Added a third temporary diagnostic (TEMP-DIAG overlap in updatePuppetSkinning, gated on bodyhairkochuru, one-shot) that computes every triangle's animated (post-skinning, not bind-pose) 2D bounding box and flags pairs that overlap on screen despite sampling UV regions more than 0.15 apart. Found a clean, repeatable cluster: low-index triangles (~10-21, UV centroid around (0.40-0.42, 0.30-0.34) - the bow) consistently overlap a distinct group of higher-index triangles (~1180-1290, UV centroid around (0.70-0.72, 0.18-0.25)). Cropped that second UV region out of the decoded source texture (bodyhair_region2.png) - it's a completely different element, a thin pale flyaway hair-wisp strand, nowhere near the bow in the texture atlas. Checked the workshop item's own preview.gif (frame crops in preview_head_*.png, upscaled 6x) - the bow renders perfectly clean there, no visible wisp bleeding into it, confirming this is a genuine rendering discrepancy versus the ground truth, not user misinterpretation.

Traced why the two stay overlapped: the bow's vertices are 100% weighted to bone 0 (the puppet root); the wisp's vertices are weighted 85-15-ish between bones 29 and 30 (parent chain 0 -> 29 -> 30). Bind-pose positions for representative vertices from each group are already only ~15-30 units apart (e.g. bow vertex 468 at (-317.2, 384.7) vs. wisp vertex 43 at (-304.0, 411.6)), so nothing in animation ever needs to separate them by much for them to stay visually overlapping. Added a fourth diagnostic logging bone 29/30's animated transform every frame instead of once (temporarily removed the one-shot guard for just these two bones) across ~150 frames of real playback: bone 29 never moves at all - position and rotation are bit-identical to bind pose on every single sampled frame; bone 30 only ever wobbles by a degree or two. Went one level deeper and dumped bone 29's raw keyframe track straight from the MDLA section, independent of the per-frame sampling code path (TEMP-DIAG bone29 raw track, in parsePuppetAnimationClips): confirmed directly from the file bytes, not just the runtime interpolation - rotZ=[0,0] and posX=[-31.7759,-31.7759] across all 181 samples of the "Animation 1" clip. This is genuine source data, not a parsing bug - the file itself has no baked keyframe motion for this bone.

Open question, not resolved this session: is a flat, unanimated bone 29 correct, or a symptom of something else this port doesn't implement? CImage.cpp's own comment on the (deliberately unparsed) MDLS second bone array claims it "turns out to hold physics/jiggle constraint parameters (angle limits, stiffness, a target position)" for each bone - which would explain a flyaway wisp having no baked track at all if the real engine instead drives it with runtime jiggle/cloth physics that this port doesn't implement, leaving it stuck at bind pose here. But docs/rendering/MDL_FILES.md's own MDLS section documentation directly contradicts this, calling the second array simply "not needed for skinning" and "not decoded" with no mention of physics data - these two pieces of this project's own documentation disagree with each other, and a from-scratch attempt this session to hand-parse the second array's per-bone record (assuming the same tmp/type/parent/matrixBytes/matrix/name shape as the first array) failed immediately (ran off the end of the buffer looking for a null terminator that wasn't there), so the physics-data claim is neither confirmed nor refuted by this session's work. Whoever picks this up next should either find where the "physics/jiggle" claim originally came from (git blame / prior session context, if recoverable) or re-derive the second array's real layout from scratch (a numberOfBones-entry array occupying exactly mdlaOffset - (end of first array) = 3290 bytes for koshini's 39 bones, i.e. ~84 bytes/bone on average, not evenly divisible - so likely still variable-length per bone, same trap as the first array). If it does turn out to be jiggle-physics data, this could be the root cause of more than just this one wallpaper's bug (worth cross-checking against Arona's still-open, structurally-unexplained halo bug mentioned earlier in this doc).

Ahri's "hole" at the torso/arm seam (2026-08-10, not yet fixed) - unrelated to the arm-mirror bug above

Same user report as the koshini section above: /workspace/1.jpg shows a jagged, striped black/orange/blue gap in spiritblossomahribase's own torso, right where ahriarm overlaps it. Given the physical location (right at the known-buggy ahriarm attachment seam), the leading assumption going in was that this is a symptom of the same still-open mirror/flip bug documented in the "Puppet attachment/cropoffset positioning bugs" section - that assumption turned out wrong.

Reproduced headlessly (tools/headless_render.sh against 3771392318) - the striped gap shows up identically to the real screenshot, as background sky color bleeding through in alternating vertical dashes, right along spiritblossomahribase's left side around screen fraction (0.44-0.58, 0.46-0.74) of a 1920x1080 capture. Note spiritblossomahribase itself renders as a flat black silhouette under this sandbox's software (llvmpipe) rendering - see the LightingV1/llvmpipe limitation noted earlier in this doc - so only the gap's presence and shape could be checked here, not its actual color/texture content; that needs real hardware.

Ruled out, with hard evidence:

  • Not ahriarm. --disable-object 44 (ahriarm) leaves the gap pixel-identical - it's not the arm failing to cover something, the gap is in the body mesh itself, independent of the arm entirely. This invalidates the initial assumption that it's the same bug as the tracked arm-mirror issue.
  • Not the leg object either. --disable-object 698 (a separate "leg" image object layered nearby) - no change.
  • Not backface culling. spiritblossomahribase.json's material declares "cullmode": "normal" (the only puppet-with-a-hole checked so far that does - bodyhairkochuru above uses nocull), which was a strong lead: if this port's puppet triangulation ever produces inconsistent winding for a subset of triangles (e.g. from the Y-flip pipeline discussed elsewhere in this doc), backface culling would discard exactly alternating triangles and produce this kind of regular striped gap. Temporarily forced CullingMode_Normal to glDisable(GL_CULL_FACE) in CPass.cpp and rebuilt - gap unchanged. Reverted.
  • Mesh coherence diagnostic (TEMP-DIAG mesh coherence for spiritblossomahribase) shows degenerate=0 outOfRange=0 - no structurally broken triangles by that check.

Follow-up, same day: same overlapping-geometry mechanism as koshini's hair-bow bug, and the two now look like the same underlying root cause wearing different clothes. All the diagnostics built for koshini's investigation above (TEMP-DIAG vertex layout candidates, TEMP-DIAG uv/pos area ratio, TEMP-DIAG overlap) were generalized to also fire for spiritblossomahribase (search CImage.cpp for isDiagTarget - now checks both object names) rather than writing new ones from scratch.

  • Vertex layout: stride=84 wins decisively (score 0.0138 vs. runner-up 0.104, not a close call) - same story as koshini, the mesh format identification isn't the problem.
  • UV/position ratio: zero outlier triangles across all 2566 - same as koshini, no local per-triangle UV corruption.
  • Overlap scan found the same pattern: low-index triangles (tri0, tri1 - UV centroid ~(0.86, 0.51-0.55)) consistently overlap a cluster of higher-index triangles (~650-800 - UV centroid ~(0.60-0.62, 0.71-0.77)) in animated screen space. Decoded materials/spiritblossomahribase.tex (tools/decode_tex.py, straightforward this time - ARGB8888/PNG-in-tex like koshini's hair, not the DXT5 format ahriarm.tex uses) and cropped both UV regions: the first is part of her tail/water-swirl art (ahri_region1.png, includes a visible "Mx" artist signature, same one seen elsewhere on the full wallpaper); the second (ahri_region2.png) is - unexpectedly - a fully painted arm, bent at the elbow with the same white fur-trim pompom visible in the user's own screenshot, baked directly into the base body texture itself, not sourced from the separate ahriarm object/texture this doc's earlier attachment-bug section is about. So this striped hole has nothing to do with the ahriarm attachment/mirror bug after all (confirmed independently two ways now: disabling ahriarm earlier in this section didn't change it, and now the overlapping geometry is traced to two other regions of spiritblossomahribase's own texture, neither of which is ahriarm).
  • Traced the vertices: the tail/water triangles are weighted to bones 55/56/59; the base-texture-arm triangles are 100%-weighted to bone 7. Checked their animated transforms at the sampled frame: bones 55 and 59 both show rotationDeg=(0,0,0) and animatedPos bit-identical to bindLocalPos - completely static, bit-for-bit the same symptom found on koshini's bone 29 (which drives her overlapping hair-wisp). Bone 7 does move, just barely (-0.89 degrees).

This is now a cross-wallpaper pattern, not a one-off: two unrelated rigs (koshini's hair, ahri's body), two unrelated decorative sub-parts (a hair wisp, a tail/water swirl), both show a driving bone that's completely frozen at bind pose (exact zero rotation and translation delta, not just small) while sibling bones in the same rig animate normally - and in both cases, the frozen bone's mesh piece ends up visually overlapping a separate, unrelated part of the same texture atlas that the real engine (per koshini's preview.gif) evidently keeps visually separated. This raises the leading hypothesis from "maybe a fluke in one file" to "probably a real, general gap in this port" - most likely something the real Wallpaper Engine drives with runtime jiggle/secondary/cloth physics that this port doesn't implement at all, leaving physics-only bones stuck at bind pose instead of settling away from other geometry.

This is not confirmed - it's the leading hypothesis, with a genuine loose end. CImage.cpp's own long-standing comment on the (deliberately unparsed) MDLS second bone array claims it "turns out to hold physics/jiggle constraint parameters (angle limits, stiffness, a target position)" - but docs/rendering/MDL_FILES.md's own documentation of the same section flatly contradicts this, calling it merely "not needed for skinning" with no mention of physics data. A from-scratch attempt this session to hand-parse the second array for koshini's rig (assuming the same per-bone record shape as the first array: tmp/type/parent/matrixBytes/matrix/name) failed immediately - ran off the end of the buffer looking for a null terminator that wasn't there in that position, so it's evidently a different, not-yet-reverse-engineered layout. This project's own documentation disagreeing with its own code comment about the same file section is itself worth resolving, independently of this bug.

Concrete next step for whoever picks this up: re-derive the MDLS second bone array's real byte layout (known total size: mdlaOffset minus the end of the first array - 20649 bytes for ahri's 73 bones, 3290 bytes for koshini's 39 bones, in case a fixed-size-per-bone guess needs a target to check division against) and check whether the "frozen" bones (koshini's 29/30, ahri's 55/56/59) have distinctly different values there than bones that do animate normally (like koshini's bone 30, or ahri's bone 7) - if so, that confirms the physics-data theory and gives a concrete data source for eventually implementing minimal jiggle simulation. If the second array turns out unrelated, the "frozen bone" pattern still needs an explanation - the next thing to check would be whether these specific bones' baked keyframe tracks in MDLA have a different internal structure (e.g. fewer real samples, a different flag) than normally-animating bones in the same file, rather than assuming they're uniformly formatted.

Follow-up, same day: cracked the second array's byte layout (a real, reproducible structural finding) - but it does not obviously explain the frozen bones. Worked this out against koshini's file directly in a scratch Python script (not added to the C++ tree - this section documents the finding rather than shipping unverified parsing code):

per-bone record (repeats boneCount times, no per-record separator):
    FLOAT  f0, f1, f2       // 3 floats, meaning unconfirmed - f1 and f2 are very often
                            // (not always) close to each other, unclear if that's structural
                            // or coincidental for these two puppets' particular rigs
    FLOAT  matrix[16]       // same row-major/translation-in-row-3 convention as the first
                            // array's bindLocalMatrix, but visibly different values (real
                            // rotation components this time, not the near-always-identity
                            // rotation the first array's bones showed)

Verified mechanically, not just by eyeballing: starting 12 bytes into the second array (the first 12 bytes look like a one-time small header/flag, still unidentified), reading exactly boneCount (39) back-to-back 76-byte records - each one's matrix passing a strict affine-2D-in-4x4 sanity check (near-zero everywhere a rotation-only 2D transform should be zero, ~1.0 on the diagonal terms that should be) - lands exactly on a boundary where record 40 immediately fails that same check (reads all zeros). That's a strong, non-coincidental confirmation of both the 76-byte stride and the 39-record count, cross-checked independently of the already-known boneCount field.

Directly compared the "frozen" bones' second-array records against normally-animating neighbors - nothing stands out. Bone 29's record (pre=(49.29, 32.52, 32.52) trans=(-11.98, 2.59)) and bone 30's (pre=(64.63, 41.59, 41.59) trans=(-15.98, -22.82)) sit well within the same range as bones that animate completely normally in the same file (e.g. bone 28: pre=(52.32, 64.72, 52.32) trans=(-12.68, -15.62), bone 38: pre=(28.04, 37.61, 28.04) trans=(-1.62, -27.66)) - no obvious flag bit, no zeroed-out field, no outlier magnitude. If this second array really does encode which bones get physics-driven, whatever marks that isn't visible from simple side-by-side inspection of the raw numbers - would need the real engine's own interpretation (i.e. an IDA disassembly of wallpaper64.exe's MDLS-second-array reader, the technique this project has used successfully for comparable format-uncertainty problems elsewhere in this doc) to actually resolve.

One more loose end found and not chased: after the 39 matrix records, koshini's file has an additional, structurally distinct 314-byte block before MDLA begins (39 × 4-byte values in a clean arithmetic progression - 0, 256, 512, ... , 9728 - immediately followed by a different, larger-magnitude sequence that doesn't obviously continue the same pattern). Unidentified; not obviously bone-index-related to the frozen-bone question, but also not yet ruled out. Whoever revisits this section should account for it before assuming the second array is "done" at the 39-record mark.

Major correction, same day: went to the real engine's disassembly for ground truth, found the second array's real per-record format is completely different from all of the above guessing - and it's empty (unused) in both of our bug wallpapers anyway, which rules out the whole line of investigation. Used IDA Pro's headless idalib (see [[lwe_ida_headless_setup]]) against wallpaper64.exe to find the actual loader, sub_1401D9860 (string-xref'd from the "MDLS0002"/"MDLA0005" version-check constants it uses - the version comparison is numeric via a helper, not a literal string match, so this one function handles every MDLS/MDLA version including our files' MDLS0004/MDLA0006).

Reading the disassembly (not the decompiler's pseudocode, which was actively misleading here - see below) of the second-array parsing loop gives the real per-record layout: [null-terminated name][DWORD id][DWORD type][16-float matrix], repeated for a count read as a WORD (2 bytes) immediately after the first array ends - not boneCount-many entries as this doc's own code comment and the earlier sessions in this file both assumed. Each record's id gets inserted into one of two unordered_map<uint32_t, int> (keyed by the numeric id field, not the name - the parsed name is stored but apparently unused for lookup) selected by the type field (0 or 1), mapping id -> array2 record index. Traced why this matters: the per-frame skeletal pose evaluator (sub_1401846E0, found by tracing the other call sites of that same map-lookup helper) contains real quaternion SLERP code that, when a flag bit on a keyframe sample is set, redirects a bone's reference transform through this exact id-map instead of using the literal bone index from its own keyframe data - i.e. a genuine constraint/aim-style redirect mechanism, not the simple bone-physics-impulse system found earlier (which is real too - applyBonePhysicsImpulse/resetBonePhysicsSimulation genuinely exist as engine-level, scripting-exposed bone physics, confirmed via decompiling both native functions - but is a separate, unrelated system operating on its own 52-byte-per-bone runtime state, not on this second array at all).

This looked like exactly the missing piece - until directly checking the count field in both actual files: it's zero in both. struct.unpack_from('<H', data, end_of_array1) reads 0 for both koshini's and ahri's puppet .mdl files. Neither wallpaper's rig uses this mechanism at all - there is no constraint/redirect data to be missing here. This retroactively invalidates the entire multi-hour byte-archaeology stream two sections up (the "cracked the second array's byte layout" [3 floats][matrix], 76-byte-stride finding, the bone-record comparison, the trailing-314-byte table) - that was pattern-matching against unused padding/reserved bytes past the (empty) array2, not real per-bone data. Floating-point reinterpretation of padding can pass a loose "does this look like a plausible rotation matrix" sanity check purely by chance, which is exactly what happened; the mistake was trusting a raw-bytes heuristic instead of verifying against the disassembly's real record boundaries from the start. Lesson for next time: when ground-truth disassembly access is available, use it before extensive black-box byte-pattern guessing, not after - it would have saved most of that earlier detour here.

Where this leaves the "frozen bone" mystery: genuinely still open, and the most likely remaining explanations have shifted. With the second-array/redirect theory eliminated, a completely flat baked track (exact zero rotation, position identical to bind pose, every sampled frame) for koshini's bone 29 and ahri's bones 55/59 might just be correct, ordinary data - not every bone needs to move, and a subtle background wisp or tail-tip bone having near-zero authored motion isn't inherently suspicious on its own.

Follow-up, same session: proved analytically that a skinning-math bug can't be the explanation either, given that data. Linear blend skinning is skinned = sum(weight_i * worldAnimated[bone_i] * inverseBindWorld[bone_i] * bindPos). When a contributing bone's worldAnimated equals its own bind-pose world transform exactly (which is what "frozen bone, zero delta from bind pose at every sampled frame" directly means, and was confirmed from the raw file bytes, not runtime behavior), worldAnimated[bone_i] * inverseBindWorld[bone_i] collapses to the identity matrix for that bone regardless of parent-chain depth or blend-weight split - and a weighted sum of identical (identity) matrices is still identity. So koshini's wisp vertices (weighted across bones 29/30, both frozen) and ahri's tail-swirl vertices (weighted across bones 55/56/59, also frozen) are mathematically guaranteed to render at exactly their bind-pose position, with zero room for a multi-bone-composition bug to introduce error - this rules out composeBoneWorldTransforms/updatePuppetSkinning's LBS implementation as a suspect for these specific cases, cleanly. Combined with the already-ruled-out effects, filtering, and mip-selection theories, this really does point back at the bind-pose vertex data itself (position and/or UV, read once at load time, never touched again) being the place a real bug would have to live, if there is one - and specifically a bug that would need to affect a whole spatially-coherent group of vertices uniformly (preserving each triangle's own internal position/UV proportionality, which is exactly what the earlier TEMP-DIAG uv/pos area ratio diagnostic checks and found no outliers for) rather than any single vertex in isolation.

Went back into the disassembly to check the real engine's per-vertex parsing directly, rather than only its bone/animation parsing - ran out of session time before reaching it. Confirmed sub_1401D9860 (the same giant function already explored for MDLS/MDLA/MDAT) is in fact the entire MDL file parser - the puppet-loading orchestrator (sub_1401894C0) hands it the whole raw file buffer from byte 0, not a pre-sliced bones-only region as earlier assumed. Its very first real statements (after ~650 lines of Hex-Rays variable declarations, this function is enormous - 14KB of code, hundreds of locals) read what looks like a version-tagged header followed by a null-terminated name table (a loop reading and hashing strings into a map, guarded by a nonzero count field) - structurally similar to TextureMap/material-name-list parsing rather than raw vertex data, suggesting the actual per-vertex position/UV/blend-weight loop (the thing that would need checking against this port's readPuppetMeshData) sits somewhat further into the function, past this name table, and wasn't reached before this session's time on the IDA angle ran out.

Concrete next step, if this thread gets picked back up: keep walking sub_1401D9860 forward from the name-table loop (ends somewhere before its first while loop's exit, look for the next place the code starts reading in large fixed-size chunks with a stride multiply - the MDLV0023 vertex block's real position/UV/blend-index/blend-weight layout should show up as exactly that pattern, the same way the MDLS bone array did) and diff its field order/offsets directly against readPuppetMeshData's assumptions (uvOffset = stride - 8, blendIndicesOffset = stride - 40, blendWeightsOffset = stride - 24) in CImage.cpp. If those offsets don't match what the disassembly shows, that's a real, fixable bug and a very different (and much safer, since it's directly checkable against ground truth rather than inferred) fix than anything guessed at earlier in this investigation. If they do match, the bind-pose-data theory is exhausted too, and the remaining candidate reverts to rendering/compositing (overlapping-triangle draw order within a single translucent-blended draw call, depth test disabled) as noted above - which doesn't need any more .mdl format work, just comparing this port's per-triangle blend behavior against the real engine's shader/blend state for the same case.

Follow-up, same session: chased the name-table loop further, found and ruled out a promising multi-submesh theory, then found the loop isn't even vertex/submesh data at all. After the name-table loop (an outer while gated by a count read as the third DWORD in the file, right after the MDLV0023 magic - at absolute file offset 17), each iteration parses what looked exactly like a per-submesh vertex-format declaration: a bitmask (built via SSE popcount against xmmword_140369CE0/_D50/_D60 lookup tables - classic "count bits set across up to 3 flag DWORDs" pattern) plus two more length-prefixed blobs, finalized via a vector-push helper (sub_1401DFFA0/sub_1401A7800).

This looked extremely promising: if a puppet mesh can have multiple submeshes each with their own explicit vertex-format flags (not a single guessed stride for the whole mesh, which is exactly what this port's resolvePuppetVertexLayout assumes), then a small decorative sub-part using a genuinely different per-vertex layout than the dominant hair-strand geometry would explain everything: why it's isolated to one small region, why per-triangle UV/position ratios still look internally consistent (each submesh's own data is self-consistent, just potentially mis-anchored if forced through the wrong globally-guessed stride), why nothing else moved it. Checked the actual count directly in koshini's file (struct.unpack_from('<I', data, 17)): it's 1. Single submesh/format-declaration, no multi-format theory to chase - directly refutes it before writing any code.

Kept reading past that loop to see what comes immediately after (in case there's still useful vertex-adjacent data nearby): it's a bounding-box min/max accumulation over a different, already-parsed vector with a 128-byte stride - matching the array2 runtime record size (name+id+type+matrix) found and subsequently ruled out (empty, count 0) in the section above, not vertex data. So this stretch of the function is still finishing up array2-adjacent bookkeeping (likely a broad-phase spatial bound for whatever array2 would be used for, harmless no-op here since array2 is empty), not the MDLV vertex buffer parser at all. The real per-vertex loop (position/UV/blend-weight, the thing that actually needs checking against readPuppetMeshData) is still further into this 14KB function and was not reached this session.

Honest status after all of this: every specific, checkable hypothesis this investigation formed - vertex-layout/stride misdetection, per-triangle UV distortion, mip/anisotropic filtering, every effect shader, backface culling (for ahri), the second-array id-redirect system, multi-bone LBS composition math, and now multi-submesh vertex formats - has been tested against real evidence (either this port's own diagnostics or the real engine's disassembly) and ruled out. What's left unverified is only the exact byte-for-byte position/UV/blend-weight field layout within the single, confirmed-correct-stride vertex block, which requires continuing to trace further into an unusually large (14KB, hundreds of Hex-Rays locals, heavy register/stack reuse) decompiled function with no natural landmarks left to search for. This is a legitimately hard reverse-engineering problem at this point, not a quick lookup - continuing productively would benefit from a fresh session/approach (e.g. anchoring the search differently, such as working backward from a known-good field like the blend-weight-sums-to-1.0 property already confirmed correct, rather than forward from the file header) rather than more of the same linear disassembly read-through.

Fixed by extending the outgoing project's lifetime past any use of the old wallpaper instead of freeing it at the point of failure: checkHotswapRequest() now collects outgoing projects into a vector scoped to the whole function (survives past its own applyAudioPolicy() call); advancePlaylist() (called once per screen, independently) pushes into a new WallpaperApplication::m_retiredProjects member instead, deliberately never pruned (a small, rare, app-lifetime retention is a much smaller cost than the use-after-free it replaces). Not reproducible in the sandbox (needs a genuinely broken/incomplete workshop asset mid-hotswap to trigger), so verified by code inspection against the crash dump's exact stack trace rather than a live repro - not yet independently re-confirmed crash-free on the user's hardware, but the "half looks like lava" report above didn't recur after this landed despite the user hotswapping heavily, which is a good sign.

Torn textures follow-up: MDLV vertex layout confirmed correct via disassembly ground truth; investigation redirected (2026-08-11)

Picked the koshini/ahri thread back up per the previous session's own suggested anchor: instead of reading sub_1401D9860 linearly from the name-table loop forward, worked outward from the fields the port's own code already leans on (uvOffset = stride - 8, blendIndicesOffset = stride - 40, blendWeightsOffset = stride - 24), and first re-derived the exact file-byte offsets those decompiled reads correspond to, since the earlier session's own offset math for "absolute file offset 17" turned out to be built on a wrong assumption about sub_140085700.

Fixed a cursor-tracking mistake from the earlier session first. Disassembled sub_140085700 directly (Hex-Rays failed to decompile it - it's a small enough function to read straight off the instruction listing): it's a NUL-terminated string reader that both returns a pointer to the string and advances the shared cursor field past the terminator. Confirmed against the raw bytes that "MDLV0023" is itself NUL-terminated in the file (byte 8 is 0x00) - so the version-string read consumes 9 bytes, not 8, and every absolute offset in the earlier session's byte-archaeology (built on an 8-byte assumption) was off by one. Re-derived from scratch: offset 9 = an unidentified per-submesh DWORD (0x1800009 for koshini), offset 13 = the inner loop's string count (always 1 in both files - a single material-name string per submesh), offset 17 = the outer submesh-count DWORD (confirmed 1 in both files, matching the earlier session's independently-verified value at that exact offset - a useful cross-check that the new math is right even though the reasoning that produced it was wrong before).

Traced the full per-submesh header field-by-field against real file bytes and found the actual vertex-format flags DWORD. After the material-name string ("materials/bodyhairkochuru.json" for koshini, "materials/spiritblossomahribase.json" for ahri), the record reads: a flags DWORD (v26, gating a conditional extra field on bit 2, unset in both files), then - because both files declare MDLV0023 and the decompiled code branches on v7 >= 17 (v7 is the parsed version number, 23 here) - six more DWORDs read as floats (v28..v33), all exactly 0.0 in both files (looks like a per-submesh bounding box field that just isn't populated for these two meshes), and finally one more DWORD (v34) that feeds directly into the SSE popcount block already found last session (xmmword_140369CE0 masks / xmmword_140369D50 sizes): this is the real per-vertex format flags word, at file offset 80 for koshini (0x180000F) and file offset 86 for ahri (0x181000E).

Computed the real stride from these flags against the actual table bytes, dumped straight out of the IDB, and it matches this port's own guessed strides exactly. xmmword_140369CE0 is a 25-entry table of single-bit masks (bits 0-24, in a fixed non-monotonic order); xmmword_140369D50 is the matching 25-entry byte-size table. (xmmword_140369CF0/_D60, referenced alongside them in the SSE loop, are provably just the same two tables read 4 elements shifted - confirmed by diffing the dumped bytes - so the whole unrolled SIMD dance reduces to one simple loop: for each of the 25 (mask, size) pairs, if the flags word has that bit set, add its size.) Summing the matched entries gives 80 for koshini's flags and 84 for ahri's flags - exact matches, independently derived from the real engine's own constant data, against the strides resolvePuppetVertexLayout's coherence heuristic already landed on for both files.

This also gives the real field order, and it matches what this port already assumes. The table's index order (not bit-value order) is what determines field order in the file, since each set flag's size just accumulates in table order. For koshini, the matched indices are 0, 2, 3, 4, 5, 6 (sizes 12, 12, 16, 16, 16, 8 - table index 1 isn't set); for ahri, they're 1, 2, 3, 4, 5, 6 (sizes 16, 12, 16, 16, 16, 8 - index 0 isn't set, index 1 is instead). Indices 0 and 1 are evidently two alternative position encodings (12-byte vs. 16-byte, mutually exclusive - confirmed by reading ahri's actual vertex bytes: the 16-byte variant is just [x, y, z=0, w=0], so reading only the first three floats, which is all this port ever does, is still correct for both variants). Indices 2-6 line up size-for-size with this port's existing hardcoded assumption of normal(12) + tangent(16) + blendindices(16) + blendweight(16) + uv(8), landing blend indices at stride-40, blend weights at stride-24, and UV at stride-8 - exactly readPuppetMeshData's current offsets. Cross-checked against the port's own already-verified fact (blend weights sum to ~1.0 at these exact offsets) which independently confirms the indices/weights aren't swapped.

Conclusion: the "concrete next step" flagged at the end of the previous session is closed out, negatively. The MDLV vertex field layout is now confirmed correct against real engine disassembly, not just self-consistency heuristics - readPuppetMeshData/resolvePuppetVertexLayout are not the bug. The two temporary diagnostics from the earlier sessions (TEMP-DIAG vertex layout candidates, TEMP-DIAG uv/pos area ratio) can be deleted next time someone's in CImage.cpp for this area; they were confidence-building scaffolding for a hypothesis that's now settled, not the bow/wisp bug.

Investigated a new lead - the real engine's actual bone-physics API - and found it doesn't explain this bug either, but the reasoning why is itself useful. Found the scripting-exposed native handlers for applyBonePhysicsImpulse and resetBonePhysicsSimulation (registration site sub_140194BF0, handlers sub_1401946C0/sub_1401949A0) and decompiled both. They operate on a 52-byte-per-bone runtime state array (offsets +40/+44/+48 hold an accumulated force vector, per the impulse handler's math) - but the array is only ever touched in response to an explicit script call naming a bone; there's no automatic per-frame simulation visible in either function, and nothing here reads gravity, wind, or any spring/stiffness constant. Checked whether either wallpaper's own script could be invoking this: koshini's scripts/ directory has only camera_paths_425.json (a camera path, not a general-purpose script), and neither wallpaper's scene.json "general" section wiring (windenabled is false for koshini; gravitystrength/gravitydirection are set but this port doesn't apply them to puppets - confirmed by grep, gravity only exists in CParticle.cpp) shows any sign of being consumed by bones specifically. Combined with the already-established fact that the MDLS second array (the only per-bone data blob that could plausibly carry physics parameters) has zero records in both files - there is no physics configuration data anywhere in either of these two files for the real engine to act on even if this port implemented the API.

This flips the framing of the "frozen bone" finding from earlier sessions. If neither file has physics data nor a script driving these bones, the real engine has no more reason to move koshini's bone 29 or ahri's bones 55/56/59 than this port does - the frozen-at-bind-pose behavior is very likely correct, not a missing-feature gap. That means the visual divergence from the real engine (clean in preview.gif, torn/holed in this port) most likely isn't about bone motion at all, and the investigation should stop looking for a reason the geometry should be in a different place. The remaining live theory is the one flagged as a fallback several sessions ago and never actually chased: how the real engine composites two genuinely screen-space-overlapping, translucent, depth-test-disabled triangle groups from the same draw call, versus how this port's rasterizer/blend state handles the identical case. That's a render-state question (blend equation, any per-triangle depth-adjacent trick, or a sort step this port's single glDrawElements call doesn't do), not a .mdl-parsing one - a materially different place to look than anything tried in this file so far.

Same session, continued: chased the render-state theory and ruled out every angle checkable without disassembling the actual D3D draw path.

  • Parent-bone motion propagation, checked and ruled out. Before trusting the earlier sessions' "frozen bone" conclusion at face value, re-examined whether it only proves bone 29's local track is static - which wouldn't by itself mean bone 29's world position is static, if its parent (bone 0, the puppet root) is animated and this port's world-transform composition failed to fold that motion in for some bones. Read composeBoneWorldTransforms/updatePuppetSkinning (CImage.cpp:445, :1420) - it's a generic recursive composition: every bone's animatedLocal is built from its own track (falling back to bind pose only for that specific bone if it has none), then composed through the real parents[] chain, with no special-casing that could strand a child bone at a stale parent transform. Confirmed live via the existing TEMP-DIAG bone anim logging (ran tools/headless_render.sh against koshini for real): bone 0 has hasTrack=1, bone 29 stays exactly rotationDeg=(0,0,0) every sampled frame (matches the earlier finding), and bone 30 does move a little (0.6-2.1 degrees, oscillating) - consistent with what's already documented, no propagation bug found.

  • Missing/unplayed second animation clip, checked and ruled out. parsePuppetAnimationClips already parses every clip the MDLA section declares (clipCount, not hardcoded to one), so if the file had a second clip - e.g. an idle/secondary motion this port simply never selects - that would be a real, findable bug. Parsed koshini's raw MDLA section directly in Python, independent of the port's own loader: MDLA0006 at file offset 104863, clipCount = 1. Only "Animation 1" exists in the file at all; nothing is being left unplayed.

  • GL blend/depth/cull state, checked directly against the real shipped shader source (not just this port's C++). bodyhairkochuru.json's material (blending: translucent, cullmode: nocull, depthtest: disabled, depthwrite: disabled, shader: genericimage4) maps in CPass.cpp:233-286 to standard glBlendFuncSeparate(GL_SRC_ALPHA, GL_ONE_MINUS_SRC_ALPHA, ...), glDisable(GL_DEPTH_TEST), glDisable(GL_CULL_FACE), glDepthMask(false) - textbook alpha compositing, no exotic state. Cross-checked against the actual shipped assets/shaders/genericimage4.frag/.vert (not reverse-engineered - these ship as plain GLSL/HLSL source with the game) with this material's combos ({}, i.e. no LIGHTING/REFLECTION/ALPHATOCOVERAGE/etc.): the fragment shader reduces to gl_FragColor = texSample2D(g_Texture0, v_TexCoord) * g_Color4 with no discard, no alpha test, no depth-adjacent trick of any kind for this combo set. GL state is a faithful match, not a divergence.

  • Got a fresh, clear repro screenshot of the actual artifact (tools/headless_render.sh, --assets-dir pointing at the real assets/ folder, no code changes) - visibly a striped/interlaced tear across the left wing's lower-left region, consistent with the wisp's thin triangles crossing the wing at a shallow angle and alpha-blending on top in narrow bands wherever they cross, exactly what "draw both groups in raw index order, no depth test" produces when they genuinely overlap on screen. Re-confirmed via the still-live TEMP-DIAG overlap diagnostic that the same triangle groups (tri10 region vs. the tri1180s cluster) still overlap in this exact run, same as prior sessions - not something that self-resolved.

Where this leaves it. Every geometry/data/state explanation reachable from this port's own source and the real engine's data files (shaders, materials, .mdl bytes) has now been checked and matches. What's left unchecked is the real engine's actual per-frame draw submission for puppet meshes - whether it issues the index buffer as-is (in which case the real engine should show the identical artifact, and the divergence must be something not yet identified) or performs some per-frame reordering/sorting before submission. That code isn't reachable the way everything else in this investigation was: wallpaper64.exe has no named DrawIndexed/DrawPrimitive-style import to search from (confirmed - idautils.Names() has no such symbols), meaning the actual D3D draw calls go through a COM/vtable interface pointer, not a plain call to a named function. Finding it would mean locating the ID3D11DeviceContext-style vtable and its DrawIndexed slot by structural/behavioral inference rather than a string or import xref - a materially bigger, more speculative next chunk of work than anything tried in this file so far, and a reasonable place for a session boundary rather than pushing further without more direction.

Same session, pushed further into the D3D/vtable hunt as directed. Real progress, but hit a genuine structural wall specific to static analysis of this binary.

  • Confirmed the renderer is real D3D11: the only import from d3d11.dll is D3D11CreateDevice itself (checked the full import table across every linked DLL) - everything else (ID3D11Device/ID3D11DeviceContext methods, including DrawIndexed) goes through the COM vtable returned by that call, exactly as suspected, not a plain symbol.
  • Found the actual call site: sub_1400FE910 (called from sub_1400E7240, the window/device-init routine) is the real D3D11CreateDevice call, storing the device/context pair at a fixed offset (+504/+512) in what's evidently a singleton Renderer-style class - Hex-Rays even auto-typed the output parameter as ID3D11DeviceContext** from the import's own prototype, confirming IDA's type library has real D3D11 vtable layouts loaded.
  • Found and fully decompiled the actual low-level draw-submission function, sub_1400C75A0, by scanning every function in the binary for the specific triplet of IASetVertexBuffers/IASetIndexBuffer/DrawIndexed vtable-offset calls (+0x90/+0x98/+0x60, the fixed ABI offsets for those ID3D11DeviceContext methods) - it's a small, clean function that just picks among Draw/DrawIndexed/DrawInstanced/DrawIndexedInstanced based on flags in its own parameter struct (instancing on/off, indexed on/off). No sorting, no reordering, nothing beyond picking the right D3D call for the flags it's given - confirmed directly from its own decompiled body, independent of anything else.
  • Chased who calls it to find where a per-object draw list might get built/sorted before submission, and hit a real dead end: the only xref to sub_1400C75A0 is a data reference from a .rdata table (off_14036C700, 6 consecutive function pointers) that looked exactly like a "RenderCommand" dispatch vtable - a very promising lead. Decompiled the apparent "constructor" using that table (sub_1400B88B0) to confirm it, and it turned out to be a lazily-initialized texture shader-resource-view cache, completely unrelated to mesh draw calls - the table is almost certainly several different, unrelated small functions/vtables that the linker happened to lay out contiguously in .rdata (COMDAT folding/section packing), not one coherent "draw command" vtable. This specific path to finding sub_1400C75A0's real caller doesn't hold up on closer inspection and needs a different anchor, not more digging on this exact lead.
  • Also traced the puppet mesh loader's own caller chain one level further (sub_1401D9860, the MDLV/MDLS/MDLA parser from earlier sessions, is called from both sub_1401894C0 - the puppet-loading orchestrator - and sub_14019BFD0) and fully mapped sub_14019BFD0: it turned out to be load-time shader-combo selection (setting SKINNING/BONECOUNT/MORPHING/MORPHING_NORMALS preprocessor combos per submesh before compiling the material's shader variant), plus an entirely unrelated second branch for meshes defined inline via "positions"/"indices"/"colors" JSON string properties (a separate, simpler custom-mesh feature, not puppets). Neither branch does any per-frame triangle reordering - both are one-time, load-time setup.

Honest bottom line. The lowest-level draw call is confirmed clean (no sorting) and the load-time mesh/material setup is confirmed clean (combo selection only) - between those two, the actual per-frame "which objects get a draw call, in what order, built from what data" logic hasn't been located, and the one concrete lead this session had for finding it (the .rdata table) turned out to be a false trail once decompiled. Finding the real per-frame render loop from here would mean identifying the specific subclass whose vtable slot 3 really is sub_1400C75A0 (there are clearly multiple unrelated classes sharing that same address purely by linker coincidence) - a proper vtable/RTTI-driven approach rather than address-adjacency guessing, which is a slower, more careful piece of work than anything else tried today.

A fundamentally better next step than more static analysis: a live capture of the real engine. This sandbox already has RenderDoc and apitrace confirmed working (per this repo's own environment notes) - a real capture of wallpaper64.exe actually rendering koshini's hair would show the actual draw call sequence, vertex data, and GPU state directly, replacing every remaining guess in this investigation with ground truth in one step. The blocker: running the real Windows binary needs Wine, which isn't installed in this sandbox (checked - wine/wine64 aren't on PATH). Installing it would be a real, if fairly standard, environment change - worth doing but flagged here rather than done unilaterally this session.

Live-capture attempt (2026-08-11): got the real engine running and tracing, but hit a genuine environment wall before reaching the puppet draw call - not a code finding, a Wine limitation

Picked up the live-capture idea flagged at the end of the previous entry. Wine turned out to already be present in the sandbox image (wine64 10.0 via libwine/wine64 packages, just not symlinked onto PATH - the real binary is at /usr/lib/wine/wine64) alongside cabextract, so no environment install was needed after all.

Isolation, done properly before touching anything GUI-capable. Per [[lwe_ida_headless_setup]]'s and [[feedback_real_display_caution]]'s standing caution about this sandbox's DISPLAY=:0 being the user's real physical monitor: started a fresh Xvfb :99 -screen 0 1280x720x24, then verified isolation by comparing xdpyinfo dimensions on both displays (:0 = 5120x1440, the real monitor; :99 = 1280x720, the virtual one) rather than just trusting the display number - a substantive check, not a cosmetic one, given the earlier documented incident where a differently-numbered display silently reconnected to the real Wayland session anyway. All Wine/apitrace work below ran with DISPLAY=:99 exclusively.

Got the real engine to actually launch and render. wineboot --init against a fresh WINEPREFIX came up clean (only the usual first-boot noise - missing bthusb/mountmgr services, common-controls manifest warnings). Confirmed real GPU passthrough is live even under Wine (WARNING: radv is not a conformant Vulkan implementation at boot - the real AMD RADV driver, not a stub). Launching wallpaper64.exe -control openWallpaper -file <koshini's real scene.pkg path> -monitor 0 -playInWindow -silent (found via strings on the exe - no -control/-file/-playInWindow docs exist anywhere, this was reconstructed from the binary's own argument-parsing strings) hit one real blocker and one dead end each dismissed/worked around:

  • A modal wine-native "Process Error" dialog ("Failed to create process: bin/ui32.exe ... code 731") appears on every launch. Root cause: this sandbox's wine64 build has no 32-bit/WoW64 support at all (confirmed independently at wineboot time - syswow64\rundll32.exe and syswow64\ntdll.dll both fail to load with c0000135), and ui32.exe (the settings/browse-UI companion process the real engine's own docs describe as architecturally separate from the D3D renderer, see the top of this file) is a genuine 32-bit binary. Dismissed the dialog each time via a synthetic click (python3-xlib + the XTEST extension, since xdotool isn't installed and there's no root to add it) - safe to do since it's on the verified-isolated :99 display, not :0.
  • Without -playInWindow, the process never produces a visible/mappable render surface at all (just the same error dialog forever, no DXGI window) - some other code path -playInWindow skips is apparently required to get a real swapchain up in this non-desktop-shell context. With it, a real "DXGI device window" and a 1284x698 "Wallpaper Pop-out" window both appear (xwininfo -tree), and D3D11CreateDevice-derived rendering genuinely starts.

Captured a real, valid apitrace GL trace of the actual engine - confirmed by inspection, not just file size. apitrace trace -a gl -- wine64 wallpaper64.exe ... intercepts the real libGL.so.1 calls Wine's wined3d backend makes when translating D3D11 to GL (this sandbox's Wine has no DXVK/vkd3d, so D3D11 goes through wined3d's GL translation - meaning apitrace's plain GL-level capture is exactly the right tool, no D3D-specific capture needed). apitrace dump on the resulting .trace file confirms it's real and coherent: genuine glXChooseVisual/glXCreateContext/D3D11CreateDevice-driven GL state setup, a real per-frame glClear/glDraw*/glBlitFramebuffer/glXSwapBuffers loop, real shader compilation (glCompileShader/glLinkProgram), and real texture uploads (glTexImage2D, including actual cubemap faces) - this is genuinely the real engine's real GL call stream, not noise.

Where it stalled, and why - a real, diagnosable environment limitation, not a rabbit hole guessed at. Let it run for a total of ~13 minutes of wall-clock capture time (two watch windows, 480s then a further 900s budget) specifically watching for the first glDrawElements/glDrawRangeElements call - the actual indexed mesh draw that koshini's puppet geometry would use (the loading-spinner UI only ever issues glDrawArrays(GL_TRIANGLE_STRIP, count=4) - simple quads, confirmed by grep, never the mesh itself). Across the whole capture: zero indexed draw calls, ever. Diagnosed this as a genuine stall rather than just slowness by checking /proc/<pid>/io: read_bytes (real block I/O, not counting cache/pipe traffic) sat at 244 KB total against a 33 MB scene.pkg - i.e. only the file's header/TOC was ever read, the bulk asset data was never touched. Texture-upload and shader-link counts in the trace (28 and 1 respectively) were also completely static for the entire second watch window despite the trace file continuing to grow - meaning all the growth was the spinner looping forever, not new content loading. Per-thread /proc inspection confirmed this wasn't a hard deadlock (the wined3d_cs thread and a real 32-wide llvmpipe-* worker pool were both still burning real, if low-duty-cycle, CPU throughout - genuinely still doing GL work each check), just work that was never going to reach the puppet mesh.

Initial theory floated here - that the openWallpaper file-load path is gated behind a handshake with ui32.exe - was checked against the disassembly immediately afterward and turned out wrong. See the following section for the correction; left this paragraph in place, struck through in spirit, so the reasoning trail stays honest rather than silently rewritten.

Checked for an escape hatch before stopping: none found. strings-searched the exe for any -no*/-skip*/-standalone/-headless-style flag that might suppress the ui32.exe companion spawn - nothing plausible turned up (-nowallpapers exists but reads as unrelated, likely a "don't auto-load the last wallpaper list" flag, not a UI-suppression one).

Honest bottom line at the time: stopping here, per this session's own instructions not to grind indefinitely on a Wine compatibility problem distinct from the original rendering bug. The live-capture approach is sound and got further than expected (a real engine process, really rendering, really traced at the GL level). The captured (but mesh-free) trace and dumps from this session are session-scratch artifacts, not committed anywhere in the tree.

Correction, same day: the ui32.exe theory was checked against the disassembly and disproven - the real stall cause is still open

Added dpkg --add-architecture i386 + wine32:i386 to /workspace/Dockerfile on the strength of the theory above, before actually verifying it against the binary. Prompted to double-check ("is ui32.exe really necessary?" - a fair challenge, since Steam's real 32-bit/64-bit launch prompt is obviously choosing between wallpaper32.exe/wallpaper64.exe, the main renderer, not the UI companion), went back to IDA rather than defending the guess.

Confirmed: ui32.exe really is the only settings/browse/editor UI binary this version ships - no ui64.exe exists anywhere in the install. So on real Windows, choosing "64-bit" at Steam launch does genuinely still spawn the 32-bit ui32.exe via WoW64, transparently and unremarkably - that part of the original reasoning holds.

But traced every xref to the "bin/ui32.exe" string and found it disproves the gating theory. All five call sites funnel through one helper, sub_14006E1E0, called with a mode string identifying which UI panel to open: WPEhandlerBrowseWallpapers, WPEhandlerSettings, WPEhandlerEditor, plus one startup-time call with no mode name. sub_14006E1E0's own decompiled body is simple: attempt process creation, and on failure just MessageBoxW(nullptr, ..., L"Process Error", 0) and return - this is exactly the dialog observed live. Nothing about koshini's scene, scene.pkg, or asset loading appears anywhere in this function or its caller's handling of its return value.

Decisive evidence: we already got past this exact dialog in the live capture, and the stall happened anyway. The captured session dismissed the "Process Error" dialog once per launch, and after dismissal the DXGI swapchain, real wined3d/GL rendering, and 28 texture uploads all proceeded normally for the following ~13 minutes - they just never reached an indexed mesh draw call. If ui32.exe were genuinely gating the file-load handshake, dismissing its failure dialog should have unblocked (or explicitly failed) the load; instead the load stayed stuck at ~244KB read of the 33MB file regardless, well after the dialog was gone. This rules the ui32.exe theory out cleanly rather than leaving it merely unconfirmed.

Where this leaves things: the real cause of the scene.pkg read stalling almost immediately is still unidentified. The ui32.exe/WoW64 angle is a real, separate, and probably still worth having (Debian's wine32:i386 package is reasonable general Wine hygiene for this app, since e.g. webwallpaper32.exe/wallpaperservice32.exe are also 32-bit-only companions per the earlier engine-layout notes at the top of this file) - but it should not be expected to fix the koshini live-capture stall specifically. The next real lead, if this gets picked back up: the scene.pkg file descriptor stayed open the entire time (per /proc/<pid>/fd) - find which thread owns it and what it's actually blocked on (a proper attach/inspect, not more guessing from the outside), or fall back to the disassembly path this investigation had otherwise exhausted (see the sections above - the still-unidentified per-frame render-loop vtable slot).

cropoffset is never added, puppet or not (2026-09-20, gura 2529364511)

Gura's 眼睛/睫毛 puppets rendered in the top right corner: resolvedOrigin=(3489,1989) instead of the scene.json origin of (2704.5,1534.5), i.e. shifted by exactly their cropoffset (784.5, 454.5). In every puppet-with-cropoffset checked in the local workshop library (gura eyes/lashes, Arona) origin - cropoffset equals the scene center exactly, meaning the editor already bakes cropoffset into origin for puppets too. The earlier "puppets need it, Arona goes off-screen without it" conclusion (section above) came from the llvmpipe sandbox, which cannot render Arona correctly, and is superseded. CImage::localTransform no longer applies cropoffset at all.

Verified on gura with a headless render (eye/lashes back on the face). Not re-verified on real hardware for the other affected puppets: Arona (2901582518, origin shifts by 773.5 px in X, may also explain her halo drifting relative to the body) and makima (3164292430, 122 px); the child puppets with tiny cropoffsets (koshini eyes, ahri arm) move by at most 8 px.

Mikasa eye rotation (3764765600), third attempt

The eye's attachment point has a static -45.09 degree rotation in the .mdl (bone angle == bind angle, it never animates). Position needs it (offset rotated by it lands on the face), orientation does not: the eye's own authored angle (+44.6) plus the bone's rotation away from rest is what orients it. Net rotation used to be ~0, which drew the lash near vertical; the preview.gif has it ~50 degrees below horizontal.

The earlier "finalAngle" attempt used about the right angle but rotated around the object origin. The eye mesh is ~66 units off that origin, so the marks left the crease. Now ResolvedTransform::meshPivotAngle (= -restAngle of each attachment point) is applied in updateScreenSpacePosition() around the puppet mesh's bounds center. Only mikasa's eye has a non-zero rest angle among koshini, ahri, Arona, asagi and mikasa's hair, so nothing else changes. Sandbox headless render matches the preview lash direction; real hardware not yet confirmed.