Skip to content

[Feature] Render the GUI on the GPU with SkiaSharp instead of Cairo #130

Description

@Zaldaryon

Problem

Every user interface in Vintage Story is rasterized on the CPU by libcairo, and the GPU only ever receives the finished bitmap. Nothing inside the interface path composites a shape, blends a pixel, or shapes a glyph itself. The rest of the client does plenty of that on the GPU: the world renderer issues GL.BlendFunc and binds framebuffer objects, and the shader set includes post-processing and upscaling. That contrast is the point. The interface is the one part of the framebuffer that never reaches a shader as anything but a finished bitmap.

The static half of that is cheap. GuiComposer.Compose() (VintagestoryApi/Client/UI/GuiComposer.cs:360) allocates one ImageSurface the size of the whole dialog, lets the static elements draw into it, optionally demultiplies alpha, and hands it to Api.Gui.LoadOrUpdateCairoTexture (:426), which ends in a GL.TexImage2D call on surface.DataPtr inside ClientPlatformWindows.LoadCairoTexture. After that Compose() early-returns while Composed is true (:362), and the dialog is drawn as one textured quad through the existing guiShaderProg path, declared in RenderAPIBase.Render2DTexture and drawn in ClientMain.

The dynamic half is what costs. VintagestoryApi/Client/UI contains 85 new ImageSurface sites, and each one is a fresh CPU surface, a full cairo raster pass, and a full RGBA upload. The trigger is a value change rather than a frame timer: the health, hunger and thirst bars when their value moves, dynamic text when its string changes, the chat caret, hover text the first time it is shown, a dropdown that opened, a slider being dragged, a number input being edited, a dialog that was resized. Every one of those events lands on the main thread, because all texture uploads have to happen on the thread that owns the single OpenGL context, and the guard on that path throws Texture uploads must happen in the main thread. We only have one OpenGL context.

Two things are worth stating precisely, because they are easy to get wrong in either direction. First, RenderInteractiveElements(float) is not where this cost lives. 43 element types in VintagestoryApi/Client/UI override it, and a brace-matched scan of every one of those method bodies finds none that allocates a cairo surface. They draw textures that were produced elsewhere. The recomposeOnRender flag (GuiComposer.cs:62, read at :721) sounds like a per-frame recompose but is only ever set from the guiScale settings watcher, in both ScreenManager and ClientMain. The per-frame cost of the interface is 43 virtual calls plus one textured quad each, and that is a draw-call question rather than a rasterization one.

Second, text is the expensive kind of dynamic change. CairoFont.SetupContext calls ctx.SelectFontFace (VintagestoryApi/Client/UI/CairoFont.cs:219) and every measurement goes through CairoFont.FontMeasuringContext (:18), a public static cairo context over a 1x1 surface. TextDrawUtil lineizes, word-wraps and autosizes against those numbers. Anything that changes its text pays font selection, glyph shaping, layout, rasterization and upload in one event.

None of this is the game doing anything wrong. It is a CPU rasterizer that has been carried forward since before the game had a GPU 2D path of its own, and today the client has no way to draw an interface without it.

Proposed change

Reimplement cairo-sharp as a recorder in front of SkiaSharp and put the interface on the GPU through it. Three objections usually kill this kind of proposal, and all three are already answered by the tree.

SkiaSharp's GL backend is already in the process. Lib/libSkiaSharp.so (9,014,008 bytes) ships next to SkiaSharp.dll, and the GL symbols are in the binary. GPU raster needs no new native dependency on Linux. SkiaSharp 3.119.1 is already pinned in three projects (Cairo/Cairo.csproj:13, sources/VintagestoryApi/VintagestoryAPI.csproj:67, Optimum.Launcher/Optimum.Launcher.csproj:16) and the launcher already pulls SkiaSharp.NativeAssets.Linux, .macOS and .Win32 for that version, so all three platforms have a known-good native set.

The client already has a GPU 2D path. RenderAPIBase.Render2DTexture draws textured quads through guiShaderProg with RgbaIn, ApplyColor and Tex2d2D, and GuiElementClip already clips with api.Render.PushScissor (VintagestoryApi/Client/UI/Elements/Impl/Misc/GuiElementClip.cs:31) rather than with cairo. Clipping and textured quad composition are not part of this work.

Font resolution is already unified. The install root ships a fonts.conf that includes /etc/fonts/fonts.conf and adds assets/game/fonts, where Vintage Story keeps its eleven bundled TTFs (Almendra, Lora and Montserrat in three or four weights each). Cairo and Skia both read fontconfig on Linux, so both see the same families and the same file list. Whether a given family resolves to the same face and the same metrics is a measurement, not a guarantee, and stage 1 is where that gets checked.

The implementation itself keeps the assembly name and every public type and signature in the Cairo namespace, and replaces the DllImport calls with managed command recording. Cairo/Cairo.csproj already owns the assembly and already declares <AssemblyName>cairo-sharp</AssemblyName> (:4), so this lands as a patch to the existing Cairo project rather than as a new subsystem.

GuiElement.ComposeElements(Context, ImageSurface) therefore keeps its exact signature, the 43 vanilla overrides keep compiling, and every mod override of it keeps compiling and keeps running. A mod already compiled against the current cairo-sharp binds to the new implementation without being recompiled, because assembly identity and the members it references are unchanged. The same applies to GuiElementTextBase.ComposeTextElements(Context, ImageSurface) (VintagestoryApi/Client/UI/Elements/Impl/Static/GuiElementTextBase.cs:47), which has 14 vanilla overrides of its own.

The ImageSurface that Compose() hands to those overrides does not go away. It becomes a recording facade rather than a rasterizer: it still has the right dimensions and the right Format, it still accepts every cairo call, and it is what the overrides draw into. Drawing into it appends display-list commands instead of pixels. Anything that genuinely needs real CPU pixels gets a real cairo surface instead, through the materialization rule below.

Recording the call sequence is what makes the fallback safe rather than hopeful. At flush time the display list executes on the GPU. If it contains an operation the GPU backend cannot honor, the whole list replays onto a real cairo context and the pixels upload as a texture, which is exactly what happens today. The worst case for an unknown cairo operation is the performance we already have, not a crash and not a missing element.

The work goes in stages, each independently useful and each behind its own flag.

Stage 1, parity harness. Before anything moves, a test renders all 43 ComposeElements overrides and all 14 ComposeTextElements overrides through both paths at GUI scale 1.0, 1.5 and 2.0, and compares cairo TextExtents and FontExtents against Skia per-glyph advances for the bundled font set. The 43 span three trees, not one: 40 sit under VintagestoryApi/Client/UI, and the other three are GuiElementMap in VSEssentials, plus GuiElementFlatList and MealstackTextComponent in VSSurvivalMod. The world map element is the reason the harness cannot be scoped to the UI folder. Golden images with a stated per-pixel tolerance, not exact equality, because cairo and Skia flatten curves differently. Every later stage is judged by this harness, and it also produces the numbers that decide whether stage 3 is worth doing on its own.

Stage 2, static compose on the GPU. Compose() stops allocating a rasterizing ImageSurface per dialog and records into a per-composer display list retained in a GPU-backed render target. The ImageSurface passed to ComposeElements becomes the recording facade described above, so no override changes shape. The existing Composed flag and ReCompose() (GuiComposer.cs:236) already mark exactly when that list has to be rebuilt, so invalidation needs no new bookkeeping. The draw stays the same single quad.

Stage 3, retain instead of re-rasterize. The 85 allocation sites get a dirty flag, so a statbar that has not moved, a chat input with no caret blink, or a hover label that has not changed keeps its display list and only re-records when its inputs actually change. This stage is caching rather than GPU work, and it needs no new backend, so it is the cheapest thing on the list to try. How much it is worth is not established yet: stage 1 is what would tell us, because that is where the recompose counts get instrumented.

Stage 4, text. Cairo stays the measurement authority. CairoFont.FontMeasuringContext and TextDrawUtil keep producing the layout, and each run is drawn at the advance cairo reported, so a rasterization difference can never move a glyph, change a wrap point or resize an autosized element. Moving measurement itself to Skia opens only after stage 1 shows per-glyph advance equality on the bundled fonts, and it gets its own flag.

Stage 5, rollout. OptimumConfig.GuiGpuRenderEnabled defaulting off, an Effective* property gated on a runtime probe that creates a Skia GL context on the game's render thread and confirms it, a switch in GuiCompositeSettings with the usual optimum-* lang keys, and a diagnostics line. This is the GreedyMeshEnabled shape exactly, including the existing convention of staying off until a repeat measurement run confirms the direction.

The main menu path (MainMenuGuiAPI, MainMenuAPI, GuiScreen) composes through the same cairo surface and takes the same treatment, since it shares the renderer.

[Implementation Note]

The compatibility surface, measured. Across the 30 mod directories in the local sample, 4 files in one family touch cairo at all. 98 files reference GuiComposer, IGuiAPI or GuiElement, but almost all of their traffic goes through the composer helpers: 69 AddStaticText, 55 AddSmallButton, 40 AddSlider, 37 AddSwitch, 22 AddShadedDialogBG, 22 AddDialogTitleBar and 19 AddButton call sites, against 15 subclasses of GuiDialog and a set of GuiInterface settings pages in the shader mods. None subclass GuiElement, and none use IconRendererDelegate.

That group matters. Those helpers live in Vintagestory.API.Client and move to the GPU as part of stage 2 without the mod author noticing, and that is where almost all mod interface traffic already is.

The direct cairo user is TerraTag, and everything it calls is a subset of the vanilla set except SetDash and IdentityMatrix. It calls new ImageSurface(Format.Argb32, w, h), new Context(surface), SelectFontFace, TextExtents, SetSourceRGBA, ShowText, SetDash, NewPath, MoveTo, LineTo, ClosePath, Fill, FillPreserve, Stroke, Paint, Scale, Save, Restore, IdentityMatrix and Dispose, then capi.Gui.LoadCairoTexture or LoadOrUpdateCairoTexture.

Taken across the vanilla interface tree and that one family, the implementation target is 68 distinct wrapper member names: 66 reached from VintagestoryApi/Client/UI and 30 from the mod. That is a third of the 204 public member names declared in Cairo/wrapper/*.cs, and it is the number the facade has to satisfy, because a recorder has to answer property reads such as Width, Height and Format as well as record method calls. Both halves come out of one scan, so the list is reproducible: collect the public member names declared in the wrapper, then collect every .Member access in the interface tree whose name is in that set. Counting method calls alone gives a smaller 38, which is the subset that has to be recorded rather than answered.

Anything that needs real CPU pixels stays on cairo and is implemented as a materialization. The scan surfaces these: ImageSurface.Data (Cairo/wrapper/ImageSurface.cs:90), ImageSurface.DataPtr (:101), the byte[] and IntPtr constructors (:53, :58), CreateForImage (:68), WriteToPng (Cairo/wrapper/Surface.cs:195), PushGroup and PopGroup (Cairo/wrapper/Context.cs:673, :685), and GroupTarget (Cairo/wrapper/Context.cs:698). Those have to keep returning genuine cairo objects, because LoadCairoTexture passes surface.DataPtr straight into GL.TexImage2D and a mod may read it too. That is what the scan found, not proof that the set is closed, so stage 1 is also where an unknown member shows up as a hard failure rather than a silent wrong pixel.

Risks. Font metrics are the real one, and a mod makes it concrete. TextExtents drives wrapping, autosize and hitbox geometry, and TerraTag already measures its own text the same way vanilla does, calling SelectFontFace and TextExtents on a context it created itself (TerraTagLabel.cs:349, TerraTagAreaLabelHudRenderer.cs:109). If the two rasterizers disagree by a fraction of a pixel per glyph, every wrapped string and every autosized element changes size. Stage 4 removes this by construction, and stage 1 measures it before anything ships.

Alpha has to stay exactly as it is. Cairo's ARGB32 is premultiplied and Compose() runs DemulAlpha() whenever the composer is not premultiplied (GuiComposer.cs:423). Under stage 2 the render target is premultiplied, DemulAlpha becomes a no-op on that path, and it stays a real operation on the materialization path so that anything reading ImageSurface.Data still gets the semantics the caller expects. Getting this backwards changes every translucent element in the game.

Blend coverage is wider than it looks. VintagestoryApi/Client/UI uses Operator.Over 835 times, and also Source 8, Clear 5, Atop 4, DestOver 1 and HardLight 1. All of them map to Skia blend modes, but each one needs a test.

Antialiased edges will not be bit-identical, because cairo and Skia flatten curves at different tolerances. That is why stage 1 uses a tolerance and not exact equality.

Mods that P/Invoke libcairo directly, or that construct a Context from a raw IntPtr handle and read Context.Handle (Cairo/wrapper/Context.cs:90, :166), bypass the facade entirely. This is rare and it has to be documented rather than assumed away.

[Acceptance Criteria]

  • All 43 ComposeElements overrides and all 14 ComposeTextElements overrides render within the agreed tolerance in GPU mode at GUI scale 1.0, 1.5 and 2.0, verified by the stage 1 harness, including the three that live outside VintagestoryApi/Client/UI.
  • Text wrapping, autosize and hitbox geometry are unchanged for the bundled font set.
  • A mod overriding ComposeElements(Context, ImageSurface) compiles and runs unchanged against the facade, including the full TerraTag call set listed above, SelectFontFace and TextExtents included.
  • Any cairo operation the GPU backend cannot honor falls back to the cairo path and produces today's pixels.
  • ImageSurface.Data, DataPtr, the byte[] and IntPtr constructors and WriteToPng still return genuine cairo surfaces, with premultiplied alpha matching what LoadCairoTexture reads today.
  • Turning the flag off restores today's code path with no configuration migration.
  • Interface composition CPU time, main-thread upload time and per-recompose surface bytes are measured before and after on at least one Windows machine and one Linux machine, with the numbers in the pull request.
  • No new native dependency on Linux. Windows and macOS verified against the SkiaSharp version the launcher already ships.

Alternatives considered

Keep cairo and only cache. Adding the stage 3 dirty flags without any of the rest touches the same 85 sites for a fraction of the work, and it needs no new backend. Whether it pays for itself is a measurement stage 1 has to produce first. It stays CPU-bound at compose time and it does not meet the goal, but it is a reasonable stopping point if stage 1 or stage 2 stalls.

Measure with cairo, rasterize with Skia, only for text. This is stage 4 of the proposal rather than an alternative, and it is the part most likely to be needed permanently.

Use a different managed 2D stack. ImageSharp has no text and no GPU raster, so text would still need Skia anyway. Silk.NET brings a renderer and windowing layer the client already has through OpenTK. SkiaSharp is already a dependency at a known version with a GL-enabled native library already deployed.

Render the whole interface into one full-screen Skia surface each frame and skip guiShaderProg. Simpler, but it forces the interface through a different blend and resolution model than the world it overlays and gives up the existing scissor and z ordering. Worth keeping as the path for elements that genuinely need to composite across several source surfaces.

Leave it alone. That is 85 CPU rasterization and upload sites on the main thread, and a font engine on the same thread that every autosized element depends on.

Scope

  • This does not change gameplay, balance, or add content.

Activity

  1. Zaldaryon commented on Oct 2, 2026

    @Zaldaryon
    CollaboratorAuthor

    Correction to the original measurements in this issue.

    The first version of the Problem section said that 33 element types allocate a Cairo surface on the RenderInteractiveElements path, and that both slider implementations allocate directly inside that method. That was wrong, and the counting method behind it was too loose. A brace-matched scan of the method bodies shows that 43 element types override RenderInteractiveElements in VintagestoryApi/Client/UI and none of them allocates inside the method. The 33 was a whole-file count that matched files containing a new ImageSurface somewhere, not on the render path.

    recomposeOnRender was also checked, since it reads like a per-frame recompose. It is only ever set from the guiScale settings watcher, in both ScreenManager and ClientMain, so it is not a per-frame cost either.

    What survives is smaller and more precise. The cost is per recompose, not per frame, and VintagestoryApi/Client/UI contains 85 new ImageSurface sites, each one a fresh CPU surface, a full cairo raster pass and a full RGBA upload, triggered by a value change rather than a frame timer. That is why stage 3 is now a standalone dirty-flag stage rather than a footnote.

    Three other numbers changed:

    • ComposeTextElements(Context, ImageSurface) (GuiElementTextBase.cs:47) has 14 vanilla overrides and is a second live mod override target. The first version of this issue did not mention it at all.
    • The target list is 55 distinct members of the cairo-sharp wrapper, counted as the union of the vanilla interface tree and the one mod family that uses cairo directly, against the wrapper's 203 public members. The first version said "about thirty" and had the wrong mix of Context members and Surface extension methods.
    • The TerraTag call set was missing ShowText, TextExtents and SelectFontFace, and listed LineWidth, LineCap, LineJoin and Operator, which that mod does not call. The missing three matter: SelectFontFace and TextExtents mean the mod measures its own text through cairo, which is the sharpest version of the font-metrics risk below.

    Also fixed: the .Add* breakdown dropped its largest bucket, since 82 of those AddCopy call sites are BlockPos.AddCopy block math rather than interface calls, and "51 GuiElement types" was replaced with two counts that can be reproduced from the tree.

  2. Zaldaryon commented on Oct 2, 2026

    @Zaldaryon
    CollaboratorAuthor

    Second correction, from an independent adversarial review of the issue body against the trees.

    Nine defects, all confirmed by re-running the check rather than by reading. Nine defects, one of which had already been corrected in the first pass.

    The ComposeElements count was scoped too narrowly. The body said 40 vanilla overrides. There are 43. 40 of them sit under VintagestoryApi/Client/UI; the other three are GuiElementMap in VSEssentials, plus GuiElementFlatList and MealstackTextComponent in VSSurvivalMod. All three are interface, so a stage 1 harness written against 40 would have shipped the world map element untested. All four occurrences now say 43, and the stage 1 text names the three.

    The implementation target list was not reproducible. The body said 55 distinct wrapper members, split 53 vanilla and 25 mod. Different reasonable scans of the same tree gave 41, 68 and 69, so the number did not mean anything. It now says 68 wrapper member names, 66 vanilla and 30 mod, against the 204 public member names declared in Cairo/wrapper/*.cs, and it states the scan that produces it: collect the public member names declared in the wrapper, then collect every .Member access in the interface tree whose name is in that set. Method calls alone give 38, and the body says so, because a recorder must answer property reads such as Width and Format, not only record calls. The claim it supports is unchanged and is the one that matters: the facade has to satisfy about a third of the wrapper, and that is a finite list.

    One citation pointed at the wrong file. The GL.TexImage2D upload that reads surface.DataPtr was cited to patches/VintagestoryLib/Vintagestory.Client.NoObf/ClientPlatformWindows.cs.patch. That patch contains no cairo code at all; the upload is unmodified vanilla in ClientPlatformWindows.LoadCairoTexture. The path is gone and the citation is now by method name, which does not leak a local mirror path.

    An overclaim in the opening sentence. "Nothing in the client composites a shape, blends a pixel, or shapes a glyph itself" is false of the client: the world renderer issues 20 GL.BlendFunc sites and binds framebuffer objects, and the shader set includes post-processing and upscaling. The contrast is the argument, so the sentence is now scoped to the interface path and says explicitly that the rest of the client does plenty of that already.

    Render2DTexture was attributed to the wrong class. It is declared in RenderAPIBase, not ClientMain. ClientMain owns guiShaderProg and draws through it, which is what the sentence was actually about.

    recomposeOnRender cited its read site, not its declaration. GuiComposer.cs:62 is the field; :721 is the if inside Compose. Every other line citation in the body is byte-exact, which made this one read as an error rather than a deliberate reference.

    The year was unsupported. "A design carried forward from 2012" had no support anywhere in the trees, the patches or the binaries. It now reads "carried forward since before the game had a GPU 2D path of its own", which is the checkable claim.

    Font matching was stated as guaranteed. "Cairo and Skia both resolve through fontconfig on Linux, so family matching can agree by construction" is not something the tree shows. They read the same configuration and therefore see the same families and the same file list; whether a family resolves to the same face and the same metrics is a measurement, and stage 1 is now named as where it happens.

    One performance claim was a guess dressed as a finding. "A measurable part of the win does not depend on the rest of the migration" had nothing behind it, and it contradicted the issue's own discipline of deferring every number to a later pull request. It now says stage 3 is the cheapest thing to try and that stage 1 is what would size it. The same sentence in Alternatives was reworded to match.

    Also tightened: the materialization list is now described as what the scan surfaces rather than as a proven closed set, which is honest, and stage 1 is named as where an unknown member shows up as a hard failure instead of a wrong pixel.

    Everything the first correction covered still holds. The 43 override count with zero allocations in the method bodies, the 85 new ImageSurface sites, the 835 Operator.Over sites with their five rare siblings, the 14 ComposeTextElements overrides, the 30 mod directories, the 98 files, the helper call-site counts, the eleven bundled TTFs, the 9,014,008 byte libSkiaSharp.so with its GL symbols, and all 24 line citations other than the two corrected above were re-measured and reproduce exactly.

  3. Zaldaryon commented on Oct 7, 2026

    @Zaldaryon
    CollaboratorAuthor

    Stage 1 result

    Stage 1 is complete as an evidence and profiling baseline. Signed commit 356745c is on feat/issue-130-gui-parity. This work does not implement the GPU renderer.

    The change adds a Linux native harness that compares libcairo with SkiaSharp's CPU raster backend, a pixel comparator with self-tests, primitive and GUI override manifests, and an API inventory for the current cairo-sharp assembly. It also adds opt-in metrics for GUI composition, text and font extent calls, and Cairo texture uploads. The metrics record counts, elapsed time, dimensions, stride/bytes, allocation versus update, and thread attribution. They do not retain GUI text. Cecil hooks make the composition, measurement, and upload measurements available in the patched game API and client assemblies.

    The ABI fixture compiles once against the installed Vintage Story cairo-sharp assembly, then runs against the candidate copy. Surface creation, drawing, text/font measurement, pixel access, and PNG output passed. This checks representative old-binary calls; it does not prove compatibility for every public API member.

    Verification

    A fresh clone of current main at ccdbbe8 bootstrapped all 122 patches. make build passed with zero warnings and errors. make check, make check-shaders, and make check-patches passed. Patch validation reported 70 source patches and 52 Cecil patches, with no pending or conflicting patches; the runtime donor check applied 22 patches and compiled the exact donors.

    All five test projects passed: Optimum.Tests 780 passed and 34 skipped, Optimum.Launcher.Tests 21 passed, Optimum.Cli.Tests 24 passed, Optimum.Bootstrap.Core.Tests 156 passed, and Optimum.Installer.Tests 78 passed. Total: 1,059 passed, 34 skipped, zero failed.

    The Cairo/Skia harness rendered 12 scene and scale combinations at GUI scales 1.0, 1.5, and 2.0. Repeated Cairo renders were byte-identical. Maximum Cairo/Skia channel deltas by scale were: geometry 2/3/2, dashed paths 52/55/41, Porter-Duff operators 1/4/1, and text 31/32/32. The harness reports these cross-rasterizer differences as diagnostics; they do not fail the functional gate. The accepted criterion is that behavior remains intact: layout and measurements, visible content, clipping, transparency, interaction, and safe fallback. Pixel equality is not required.

    The focused Pharos test opened the inventory dialog and captured a 1024×768 frame. It found 18 GUI elements, 8 interactive controls, and recorded 2 compositions, 69 texture uploads (40 allocations and 29 updates), 21 text extent measurements, and 7 font extent measurements. This verifies one live dialog and the instrumentation, not the full GUI matrix. No manual review of every client screen is claimed.

    Limits and next step

    The primitive manifest marks unrun operations explicitly. The GUI manifest inventories 57 composition overrides, but source discovery is not counted as rendering coverage. Every operation without a reviewed passing fixture stays on native Cairo. The API inventory records 81 exported types and 868 public/protected members with initial categories. Human review remains necessary before claiming full ABI compatibility or allowing any member onto a replacement backend.

    The next step is to review the API members needed by the recorder and lossless Cairo fallback, then implement that path with unvalidated calls materializing to native Cairo before execution. GPU context creation, GPU rendering, full GUI coverage, and Windows/macOS validation remain future work. Atlas results are not counted as GUI evidence because Atlas does not create the client's GUI or GL context.

  4. Zaldaryon commented on Oct 8, 2026

    @Zaldaryon
    CollaboratorAuthor

    Stage 1 merged

    PR #136 was squash-merged into main on 2026-10-08 as 72892d5. The merge contains the Cairo/Skia CPU parity harness, GUI coverage manifests, Cairo API inventory, old-binary ABI fixture, and opt-in GUI composition, text measurement, and texture upload metrics. It does not add a GPU renderer.

    Validation on the merged head passed make check, make build, make check-patches, make check-shaders, and all five test projects: 1,059 passed, 34 skipped, 0 failed. Patch validation found 70 source patches and 52 Cecil patches, with no pending patches or conflicts, and compiled the 22 runtime donors. The Cairo parity harness produced byte-identical repeated Cairo renders at scales 1.0, 1.5, and 2.0. Cairo versus Skia differences remain diagnostic. The old-binary Cairo ABI fixture passed against the installed Vintage Story 1.22.7 assembly.

    A separate Pharos v0.5.0 run against Optimum commit 356745c exercised 11 live GUI states and captured 11 non-uniform 1280x720 screenshots with zero scenario errors. Its usage trace matched all 868 API manifest signatures and recorded 112 called members. The inventory hover selected slot 0, and the map and chat were focused after input. These states are live usage evidence, not rendered coverage for every GUI override.

    The make run smoke test launched the Linux client, loaded the GUI shader, and reached the login screen. The cached session key was invalid, so this run did not enter a world. All 57 composition override rows remain not-rendered and unassessed; manual review of the API classifications and the full GUI matrix remain open. This merge completes the Stage 1 evidence baseline only. No GPU compatibility, pixel parity, or performance claim is made. The next implementation stage is the recorder with lossless native Cairo fallback.

  5. Zaldaryon commented on Oct 8, 2026

    @Zaldaryon
    CollaboratorAuthor

    Stage 2 checkpoint

    Signed commit 4c32732 is published on feat/issue-130-gpu-gui-rendering.

    This checkpoint adds a Cairo surface recorder with native Cairo materialization for unsupported operations, a runtime GPU probe and candidate texture path, GUI scale capture, and recorder/probe tests. The default setting remains off. The GUI still falls back to native Cairo when the GPU candidate rejects an operation or GL state.

    Verification

    All five unit test projects passed with 1,093 passed, 34 skipped, and zero failed. make build, patch IL generation, patch validation, compatibility checks, shader checks, patch syntax checks, and git diff --check passed. Patch validation reported 74 source patches and 52 Cecil patches, with no pending, unavailable, or conflicting patches; 22 runtime donors compiled.

    The Pharos GUI suite passed both tests. It captured 11 live states at the baseline and GUI scales 1.0, 1.5, and 2.0: 44 non-uniform frames with zero scenario errors. Three synthetic GL error and recovery cycles restored state and succeeded on the next candidate. Scale-specific inventory captures recorded six successful GPU candidates across scales 1.5 and 2.0.

    Remaining work

    The primitive manifest has 20 rendered rows and 8 partial rows. The 57 GUI override rows still have no per-row visual acceptance, with 0 of 171 row and scale comparisons completed. The inventory timing runs measured GPU-requested recomposition at 714 ms versus 591 ms with GPU off, a 21% increase in this short workload. Those timing runs had zero successful GPU candidates, so this is a fallback overhead signal. Memory retention and broader performance measurements remain open.

    This is an implementation checkpoint, not completion of the feature.

  6. Zaldaryon commented on Oct 8, 2026

    @Zaldaryon
    CollaboratorAuthor

    Stage 2 follow-up

    Signed commit aa7d8e7 is now on feat/issue-130-gpu-gui-rendering.

    The Cairo fallback fixtures now cover all 27 non-text primitive rows, leaving only text layout and wrapping partial. The parity harness has an optional font matrix that verifies Cairo resolves all 11 bundled TTF faces and captures 33 face and scale combinations. Cairo stayed deterministic. Cairo and direct-file Skia text advances and pixels still differ, so these measurements remain diagnostic and Cairo remains the layout authority.

    The GPU candidate path now rejects surfaces that already materialized to Cairo before checking for a current GL context or running the one-shot probe. Pharos verifies the native-surface rejection while preserving the GL sentinel state.

    Verification

    All five unit-test projects passed with 1,096 passed, 34 skipped, and zero failed. make patch-il, make check-patches, make check-shaders, and git diff --check passed. Patch validation reports 74 source patches and 52 Cecil patches, with no pending, unavailable, or conflicting patches; 22 runtime donors compiled.

    The complete Pharos GUI suite passed both tests. Its live trace contains 11 scenarios, 44 captures at scales 1.0, 1.5, and 2.0, 115 distinct Cairo API members, and zero scenario errors. The inventory attempted 403 GPU candidates, with 6 successes and 397 Cairo fallbacks. Synthetic tests recovered after an injected GL error and preserved GL state.

    Two short A/B runs of 10 inventory recompositions each averaged 513 ms with GPU requested and 528 ms with GPU off. Individual results varied from 498 to 568 ms, and a third GPU-requested run measured 566 ms. This is inconclusive and does not establish a performance gain.

    Remaining work

    Text wrapping, layout and hitbox parity remain open. The 57 GUI override rows still have 0 of 171 row and scale comparisons. Retained memory and broader performance measurements also remain open.

  7. Zaldaryon commented on Oct 8, 2026

    @Zaldaryon
    CollaboratorAuthor

    Stage 2 text and runtime coverage follow-up

    Signed commit bcd84e6 is published on feat/issue-130-gpu-gui-rendering. It adds nine regression cases for text wrapping, line bounds, multiline height, autosizing, and Cairo fallback output across GUI scales 1.0, 1.5, and 2.0.

    All five test projects passed with 1,105 passed, 34 skipped, and zero failed. make check, make build with zero warnings and errors, make patch-il, make check-patches with 74 source patches, 52 Cecil patches, zero pending or conflicting patches, and 22 runtime donor compilations, make check-shaders, and git diff --check passed.

    The isolated Pharos run exercised 11 GUI states and captured 44 frames across the baseline and all three scales, with 115 Cairo API members and zero scenario errors. Direct override tracing recorded 29 of the 57 manifest rows at each scale, including GuiElementMap. The other 28 rows did not run in the available states. This is runtime invocation evidence; 0 of 171 row and scale comparisons have an accepted visual tolerance, so the visual matrix remains open.

    The Radeon probe passed 32 injected-error and recovery cycles, preserving GL sentinel state and verifying texture deletion after each successful render. The inventory attempted 403 GPU candidates, accepted six, and fell back for 397. One short 10-recompose sample measured 502.1 ms with GPU requested and 516.5 ms with it off. That result is inconclusive for performance. Windows and macOS runtime checks were not available on this Linux host.

  8. Zaldaryon commented on Oct 9, 2026

    @Zaldaryon
    CollaboratorAuthor

    Signed commit fb8ae08 is published on feat/issue-130-gpu-blur, based on the previous bcd84e6 checkpoint.

    Recorded GUI surfaces with blur now keep drawing on the GPU. RecordedGpuBlur runs three horizontal/vertical RGB box-blur pairs through Skia runtime shaders, preserving alpha, channel truncation, clamped borders and the original partial-blur accumulator behavior. GPU snapshots and two reusable GPU surfaces replace the CPU raster step in the candidate path. Shader effects are cached by radius with a 16-entry limit and released with the owning context. The commit also includes the pending supported-command recorder work, texture-orientation correction, context reuse and teardown, and candidate metrics fixes. Native Cairo replay remains available on failure, including narrow partial rectangles with overlapping box windows. The experimental setting remains off by default.

    Verification on Windows, Vintage Story 1.22.7, SkiaSharp 3.119.1 and an AMD Radeon RX 6650 XT:

    • Fresh bootstrap from the official installer: 126 patches applied, zero skipped or failed. Release build passed with zero errors and 101 decompiled-code/analyzer warnings.
    • All five unit-test projects: 1,135 passed, 34 skipped, zero failed. Ten new shader fixtures provide deterministic raster coverage for CI.
    • Actual GPU blur: 120 matched-input fixtures covered full/partial blur, transparency and drawing between two blurs. Each executed 12 shader passes. Maximum RGB difference from the native blur was one level; alpha matched exactly.
    • 4,096 GPU stress cycles passed GL state restoration, injected-error fallback, thread guards and texture deletion. All 32 context recreations also rendered blur. Private process memory after full GC changed from 190.38 to 180.91 MiB; this does not measure VRAM. Separate tests passed shader-cache eviction and native replay after a GPU blur rejection.
    • Patch checks: 74 source patches and 52 Cecil patches, zero pending, unavailable or conflicting patches. All 22 runtime patches applied and exact donors compiled. Shader validation and prerequisites passed; two optional packaging tools were unavailable, and compatibility checks retained 17 baseline skips.
    • The final isolated Pharos GUI run captured 66 frames across 11 states and three scales, with zero client errors, trace gaps or element-bound mismatches. Six comparison sheets were inspected. The normal staged client also started, passed the Radeon probe and closed with exit code 0. Existing settings were restored and checked by SHA256 after the runs.

    The inventory timing run used seven alternating trials per mode and scale, three warmups and 20 recompositions per trial, without concurrent bootstrap, builds, tests or GPU stress. All 420 measured GPU candidates succeeded with zero fallback and 5,040 blur shader passes.

    GUI scale Cairo median ms/recompose GPU median ms/recompose Reduction in median time
    1.0 40.72 37.81 7.1%
    1.5 67.30 42.66 36.6%
    2.0 88.98 51.30 42.3%

    Median paired reductions were 7.1%, 34.8% and 36.1%. These results measure inventory recomposition with tracing and shared-process state; they do not establish FPS gains or a speedup for every GUI workload.

    The remaining feature work is visual acceptance of all 57 composition overrides at three scales, source-color and antialiasing parity, runtime checks on Linux/macOS and other GPU vendors, and broader dynamic-GUI/resource measurements. Whole compositions are not byte-identical to Cairo: a separate diagnostic painting random translucent colors independently reached a 28-level RGB difference after partial blur. The matched-input kernel tests do not resolve that source-color conversion difference. Captures also vary with animation, map growth, chat and hover; large scales clip at 1280x720 in both modes. This publishes the GPU blur implementation and its Windows evidence while keeping issue #130 open for the remaining renderer acceptance work.

  9. Zaldaryon commented on Oct 9, 2026

    @Zaldaryon
    CollaboratorAuthor

    Follow-up plan for keeping the complete GUI rendering path on the GPU.

    The measurable target is zero CPU pixel rasterization and zero GPU-to-CPU image transfers while composing supported interfaces. Input handling, UI logic, layout, text shaping and command submission still consume CPU time. Skia's GPU backend also performs CPU preparation. A literal interface with 0% CPU usage would require a different execution architecture and would still need host coordination.

    The current checkpoint is fb8ae08. Recorded supported drawing and the six blur passes already execute on GPU surfaces. This is a proposed next phase; the items below are not implemented by that commit.

    1. Replace the Cairo shadow drawing context with a retained command and state model. Record paths, transforms, clips, paint and source references directly, including save/restore semantics. The current recorder still uses Cairo CopyPath, text outlines and source snapshots. Inventory every call made by vanilla interfaces and mods, then implement equivalent behavior for each supported operation. Skia can execute the drawing on the GPU, although path preparation may still use the CPU.
    2. Keep surface dependencies on the GPU. Make recorded surfaces reference GPU textures or retained recordings, with explicit ownership, dirty generations, context-loss recovery and bounded caches. Compose nested surfaces, clips, gradients, image scaling, blending and blur without materializing CPU pixels. Upload external image assets when loaded or changed, and reuse their textures.
    3. Give text a dedicated retained path. Cache shaping, metrics and layout; reuse glyph geometry or atlases across recompositions. To eliminate CPU glyph rasterization as well, use a verified GPU glyph-generation path, such as rendering outlines into a GPU atlas. Choosing a GPU canvas alone does not establish that font rasterization is GPU-only. Preserve font fallback, hinting, scale behavior and existing layout measurements.
    4. Complete the operation and compatibility matrix. Pixel access through DataPtr, unsupported Cairo operations and consumers that require a CPU image currently force native materialization. Each needs an explicit implementation or a documented compatibility fallback. Achieving zero CPU rasterization for every interface, including mods, requires resolving those consumers too.
    5. Remove blocking synchronization from ordinary composition after validating ordering and resource lifetime. The candidate currently calls Submit(true) and GL.Finish(). Establish safe submission and synchronization between Skia and the game's OpenGL renderer, then measure the asynchronous path. Keep diagnostic readbacks outside ordinary composition.
    6. Reuse retained recordings and textures until content changes, updating only dirty regions where correct. Batch compatible draws and avoid rebuilding unchanged interfaces. This reduces CPU recording work alongside GPU rendering work.

    Acceptance should report CPU composition time, GPU execution time, CPU rasterizations, native fallback count, readback count, uploads and VRAM separately. For the supported path, require zero CPU rasterizations, zero readbacks and zero native fallbacks; exercise the full interface/scale matrix, dynamic updates, nested surfaces, text, partial blur and context recreation. Keep Linux/macOS, other GPU vendors and visual parity in the validation scope.

    A further project could move selected layout or geometry algorithms into compute shaders. Moving arbitrary UI and mod logic would require changing their execution model and APIs. It should have its own scope and evidence before any claim of zero CPU UI work.

    Skia runtime effects describe how shader effects participate in the drawing pipeline. Our backend must also guarantee GPU-backed targets and GPU-resident dependencies to meet the rendering target above.

  10. Zaldaryon commented on Oct 9, 2026

    @Zaldaryon
    CollaboratorAuthor

    Pushed signed commit 1383f3a to feat/issue-130-retained-ui. This extends the GPU path to retained source surfaces, internal GUI textures, bitmap transforms and SVG rendering. Immutable source generations preserve surface mutation and disposal behavior; recording can resume after upload. Queued text/statbar context handoffs now close their writers before GPU upload. SVG geometry is recorded from NanoSVG into Skia pictures without ordinary NanoSVG pixel rasterization. A doubled-alpha caret stroke found by visual comparison is fixed and covered by regression tests.

    Linux validation used Vintage Story 1.22.7, SkiaSharp 3.119.1 and a physical Radeon GPU with Mesa 25.2.8, using the normal authenticated datapath:

    • Clean archive reconstruction: 134/134 patches applied, build with zero warnings/errors, 1,139 tests passed and 34 existing skips across all five projects. Patch, shader and compatibility gates passed; compatibility retains 17 documented skips. One existing animation diagnostics test failed on the initial run, then passed both a focused rerun and the full suite; the cause is not established.
    • Inventory: 467/467 GPU candidates, zero fallback. Sixteen live UI states and 60 supplementary fixture cases at scales 1.0, 1.5 and 2.0 showed no ordinary CPU rasterization, native surface materialization or GPU readbacks. Startup still performs one diagnostic capability readback.
    • Automated visual acceptance: 1,955 matched GPU/native recording-prefix samples cover all 57 overrides at all three scales, accepting 171/171 cases. Eighteen cases produce no Cairo pixels. Every native reference was rendered twice with identical bytes. The comparator passed 11 self-checks, including rejection of missing glyphs, translations, clipping, wrong color/alpha and the original caret defect.
    • Existing Cairo assembly compatibility: 81 public types and 868 public/protected members retain their ABI; an old compiled mod fixture passed. Thirty-two injected GL-error recovery and texture-disposal cycles passed.

    Visual acceptance uses one shared antialiasing policy, not pixel identity: flat premultiplied RGBA tolerance 3, edge distance 1.5 pixels with at most 0.2% unmatched edges, alpha-mass difference at most 24 per occupied pixel, and Gaussian-filtered limits of maximum 96, content mean 8 and 99th percentile 55. Comparisons use content bounds and transparent image borders. This accepts bounded Cairo/Skia font and SVG antialiasing differences. Coverage is limited to the captured live states and fixtures, not every possible UI content or mod.

    Performance still needs work. A matched ten-recomposition Linux sample measured about 952 ms with GPU rendering versus 493 ms with CPU rendering. Increasing the Skia cache did not improve that result, so this checkpoint makes no performance gain claim. Cairo still supplies layout, shaping and native shadow state, and explicit unsupported native escapes retain the compatibility fallback. Removing that shadow, improving dirty reuse, measuring VRAM retention and validating the new implementation on Windows/macOS remain open. The issue stays open; no PR or merge is included in this checkpoint.

  11. Zaldaryon commented on Oct 9, 2026

    @Zaldaryon
    CollaboratorAuthor

    Implementation plan: reduce CPU recomposition and remove native Cairo

    Continuation point: signed commit 1383f3a, branch feat/issue-130-retained-ui. The preceding checkpoint records the Linux validation and its limits. This comment is a plan; the work below has not been implemented or benchmarked.

    The target has two independently verifiable outcomes: all supported ordinary UI rasterization/effects stay on the GPU, and the client can start and exercise that UI without loading native Cairo. Reducing main-thread CPU time is a third measured acceptance condition. Removing Cairo alone does not establish that reduction.

    What can move, and what cannot

    SkiaSharp provides GPU drawing, recorded pictures and shader effects. It does not execute existing C# GuiElement callbacks, bounds calculation, font selection or application event handling on the GPU. SKPicture records drawing commands; recording and playback still have CPU/driver costs. Shaping, line breaking and metrics can use a Cairo-free CPU implementation, with caching and immutable worker preparation. Input needs CPU-visible bounds, caret positions and hitboxes. Moving arbitrary layout to a compute shader would require a new layout engine and results synchronization, outside the ordinary SkiaSharp drawing API.

    Current cost Planned replacement Where work remains
    Bounds, wrapping and autosize Versioned layout and measurement caches; recalculate affected subtrees CPU, with safe pure calculations eligible for workers
    Native Cairo shadow and getter queries Managed drawing state and immutable path/source descriptions CPU state bookkeeping, no Cairo replay
    Text-to-outline conversion per draw Cached compatible glyph runs/outlines, then validated SKTextBlob drawing CPU shaping on a cache miss; GPU drawing, potentially CPU glyph preparation
    Rebuilding every element's commands Retained element/display-list generations and explicit invalidation CPU preparation only for changes
    Surface allocation and full redraw Reuse GPU targets, dependency images and static layers; redraw damaged regions GPU raster/composition with bounded resource management
    Many GL transitions and submissions One controlled rendering boundary, ordered batches, asynchronous submission CPU/driver submission remains necessary
    Native pixel/handle escape Cairo-free compatibility implementation where possible; explicitly classified exceptions CPU pixels only when the caller requests them

    The architecture follows Skia's GPU surface and picture model. SkSL runtime effects implement shaders, color filters and blenders; they are not a general replacement for C# layout. A GPU surface must use the correct current GL context. CPU preparation on a worker does not authorize sharing a mutable GRContext across threads.

    Libraries and functions for a Cairo-free implementation

    Candidate Use Decision and validation
    Existing SkiaSharp 3.119.1: SKPath, SKPaint, SKMatrix, SKShader, SKPictureRecorder, SKFont, SKTypeface, SKTextBlob Drawing state, geometry, metrics, retained commands and text rendering Primary renderer. Inspect the exact pinned binding and native export availability before choosing signatures. CPU-backed SKSurface is the Cairo-free software fallback when GPU rendering is unavailable.
    SkiaSharp.HarfBuzz / SKShaper or direct HarfBuzzSharp Glyph substitution and positions for complex scripts Prototype and compare with the existing text behavior. HarfBuzz shaping maps characters to positioned glyphs; it does not replace the entire UI layout engine. Match package/native versions and package Linux/Windows/macOS assets.
    Existing FreeType access, or a narrow FreeType metrics adapter Font-face identity, hinted advances, bearings and outlines if Skia metrics fail compatibility Use only when differential tests establish a need. FreeType is a font engine, not a complete layout/shaping replacement. Match load flags, hinting and rounding; do not assume identical Cairo measurements.
    Topten.RichTextKit Optional paragraph layout, bidi, line breaking, caret/selection and hit testing Research candidate, not the default migration. Its README documents HarfBuzzSharp, fallback and editing support, but reports Windows testing and untested other platforms. Verify Linux behavior and compatibility with pinned SkiaSharp first. Replacing existing wrapping can change dialog sizes and interaction.
    Current NanoSVG geometry parser plus SKPicture Existing SVG assets Retain first. NanoSVG parsing does not require Cairo. Remove native raster fallback closures when the Cairo-free renderer can represent their semantics. Cache parsed geometry by asset content and viewport.
    Svg.Skia Optional richer SVG support into SKPicture Evaluate only if the asset corpus needs unsupported SVG features. It is not necessary to remove Cairo. Compare gradients, transforms, viewport, tint and alpha/color-order behavior; pin a compatible version and audit transitive/native assets and license notices.

    Recommended first implementation: existing SkiaSharp plus a small managed compatibility layer, existing NanoSVG parsing, and a text provider prototype using Skia metrics with FreeType/HarfBuzz only where needed. A new library is not a blanket solution for Cairo's semantics. Do not replace the entire UI toolkit or upgrade the graphics backend as part of this migration without a separate measured reason.

    Proposed internal boundaries, not existing public APIs: GuiDrawingState owns double-precision matrices, path/current-point semantics, paint, source and save/restore state; IGuiTextMetricsProvider returns Cairo-compatible extents and font metrics; GuiTextRunCache owns immutable glyph IDs, positions and cluster maps; GuiLayoutSnapshot holds bounds, baselines, wrapping and input geometry; GuiDisplayList holds immutable draw operations and resource generations; GuiGpuScheduler batches those lists on the GL-owning thread. Keep these implementation details behind the existing public wrapper.

    Ordered work and acceptance gates

    • P0. Profile the cost before changing its placement. Add disabled-by-default spans/counters for bounds calculation, TextExtents/font metrics, native shadow replay, path conversion, source snapshots, recording, allocation, GPU replay, state capture/reset/restore, flush/submission and queue latency. Measure thread CPU time separately from elapsed time and waiting; process-wide CPU time alone cannot identify main-thread relief. Count allocations/GC, commands, cache hits/misses, texture/FBO creation, uploaded bytes, native calls, fallback reasons, readbacks and blocking waits. Retrieve supported GL GPU timer queries asynchronously on later frames; discard disjoint/invalid samples. Compare baseline CPU mode and current GPU mode with diagnostics/visual readbacks disabled during timing. The approximately 952 ms GPU versus 493 ms CPU ten-recomposition result is elapsed time and does not establish CPU savings or a bottleneck cause. Deliver an attributed baseline before selecting the largest optimization.
    • P1. Build managed state while retaining the native oracle in tests. Replace SurfaceRecorder.Shadow and RecordedGuiDrawing.Capture(Context, ...) getter round trips with explicit state. Record path segments once, preserve double precision until the Skia boundary, and implement transforms, current point, relative commands, NewPath/ClosePath, fill rules, save/restore, clips, dashes, cap/join/miter, source binding transforms, extend/filter, operators and antialiasing. Store source descriptions without creating a native Pattern for every draw. Preserve immutable source-prefix generation, mutation/disposal and owner-thread semantics. Native replay becomes a diagnostic oracle or temporary migration fallback, not the normal path. Gate: state/geometry differential fixtures pass, ordinary covered drawing creates no native shadow or Cairo context, and the caret retraced-stroke fix still passes. Cache stroked outlines by geometry/style/scale instead of repeatedly calling GetFillPath; benchmark before changing to an alternate stroke renderer.
    • P2. Remove Cairo from font metrics and text construction. Audit CairoFont.FontMeasuringContext, CairoFont.GetTextExtents, TextDrawUtil.Lineize, recorder TextExtents/TextPath/GlyphPath, and direct wrapper/mod entry points. Prototype providers against the exact bundled font files, face indices, size, matrix, font options and rounding. Preserve width versus advance, bearings, ascent/descent, baseline, current-point advancement, trailing spaces, wrapping thresholds, AutoBoxSize, AutoFontSize, flow paths and selection/hitbox behavior. Cairo's simple text API must not silently gain HarfBuzz ligatures/kerning that change existing measurements; separate compatibility text from explicitly shaped runs. For ShowGlyphs, retain supplied glyph IDs/positions and prove font-face mapping rather than reshaping. Cache measurements and glyph runs using font content identity, options, scale, text, direction/language/features and font-fallback generation. Start with cached compatible outlines; adopt text blobs only after text parity and cost comparison. A text blob may still cause CPU glyph generation and atlas upload on a cold miss; measure these rather than calling it zero CPU work. Gate: exact layout decisions and input geometry where currently deterministic, accepted pixels under the fixed policy, and no normal font-metric/path call into Cairo. Keep a Cairo reference only in development tests.
    • P3. Make layout and commands incremental. GuiComposer.Compose already returns early when Composed is true; preserve that fast path. For actual ReCompose, separate layout dirtiness, paint dirtiness, transform-only changes and external resource generations. Cache static element command lists and text layout; rebuild only affected dependencies. Keys include content/style/font revisions, GUI scale, constraints, theme/locale and source generations. Bounds changes propagate to dependent parent/children; asset/font reload invalidates the proper cache. Unknown mod callbacks and publicly mutable fields require explicit invalidation hooks or conservative whole-component rebuilds, not stale cached output. Never skip OnComposed, focus changes or callback side effects. Gate: repeated unchanged state causes no new measurement/path recording or target allocation; one changed field rebuilds the required subset, and bounds/focus/callback behavior matches baseline.
    • P4. Reuse GPU targets and render damage correctly. Audit TryCreateCandidateTexture and LoadOrUpdateCairoTexture: the current candidate path allocates a texture/FBO; the engine-side larger-texture reuse does not prove candidate reuse. Retain targets by dimensions/format/filter/context generation and keep dependency caches bounded. Start with static background plus dynamic overlays, then dirty rectangles. Clear the full damaged region and replay every intersecting element in original order, including elements above a changed one. Expand damage for stroke, shadows, blur support and old/new bounds; preserve clip, nonlocal operators, premultiplication and transparent padding. Full redraw is the conservative path for unbounded effects. Keep logical size separate from allocation size/UVs. Retain old source generations while referenced; do not overwrite a texture still used by a snapshot or queued render. Gate: correct overlap/clear/scroll/resize behavior, bounded GPU memory, no steady-state target churn, and measured lower CPU submission/allocation cost.
    • P5. Batch the shared-GL rendering boundary. Queue related dirty surfaces/dependencies and render them in dependency order under one safe engine/Skia state boundary where call-site semantics permit. Preserve immediate texture availability when callers require it. Audit redundant skSurface.Flush, context.Flush, state queries and ResetContext per surface; remove only with verified resource visibility and engine-state restoration. Use asynchronous submission; reserve synchronized submits/readbacks for explicit diagnostics or requested CPU pixels. Record queue length and end-to-end latency; do not move a stall to the next frame and report it as eliminated. Same-context GL ordering is preferable to a new shared context; cross-context rendering requires separate fences and a proven resource handoff. Gate: sentinel GL tests, error recovery, texture lifecycle and real-world map/item rendering still pass, with fewer transitions/submissions and no ordinary blocking completion wait.
    • P6. Move safe preparation off the main thread. Capture immutable state on the owning thread, then prepare pure layout calculations, independent text-provider instances, SVG parsing and immutable display lists on a bounded worker queue. Do not run arbitrary GuiElement/mod callbacks on workers; they can access mutable game state and APIs. Do not share the public mutable FontMeasuringContext or mutable Skia objects. Use generation cancellation/coalescing, bounded memory and stale-result rejection. Atomically commit matching visual/input snapshots on the main thread. Keep small latency-sensitive edits synchronous if necessary; do not delay caret/hitboxes independently of their visuals. All GL work stays on the current GL-owning thread initially. Gate: deterministic results, no races/use-after-dispose, prompt cancellation on dialog close or scale change, bounded backlog, and reduced main-thread CPU time without increased total CPU cost or visible input lag.
    • P7. Replace native compatibility escapes and surface backends. Classify every member in the existing API inventory, not only the 57 overrides. Implement supported image surfaces, stride/data access, Flush/MarkDirty, groups/masks, copy/path queries, image encoding, font-face/scaled-font access and software rendering through managed state/Skia. Materialize CPU pixels lazily only when requested; a GPU readback for DataPtr must be explicit, counted and followed by correct mutation invalidation. Preserve the existing assembly/type/member identity and lifetime rules for managed consumers. Audit PDF/PS/SVG and platform surface classes such as Xlib/Win32/Xcb/Glitz: preserving a signature is insufficient if behavior is missing. Implement required cases, or document a separately scoped capability limitation. Gate: old compiled consumers pass against the candidate, all inventory rows have evidence or an explicit limitation, and the Cairo-free software mode passes when GPU initialization fails.
    • P8. Remove the runtime dependency and prove it. Replace every reachable Cairo/NativeMethods.cs import (libcairo-2), font/surface factory and native replay closure. Audit donors, Cecil hooks, build/bootstrap references, packaging manifests and native dependency chains. Retain the Cairo compatibility assembly name where mods require it; its name does not imply a native Cairo dependency. Run a packaged client in an environment where Cairo loading is denied, including transitive loading; verify loaded-module/readback evidence, not just absence of a bundled DLL. Exercise startup/fonts, all UI fixtures and live states, software fallback, device reset, resource teardown and selected real mods. Remove obsolete native files only from the explicit package/build inputs. Gate: no mandatory Cairo import, no loaded native Cairo module and zero Cairo calls for the supported package path. Package any newly needed HarfBuzz/FreeType assets and notices, verifying all supported architectures and operating systems.
    • P9. Qualify performance and release readiness. Run the expanded correctness matrix and matched CPU/GPU A/B benchmarks below, then enable the new backend progressively. Retain independent configuration for GPU rendering and compatibility mode during migration. Remove the native compatibility option only after its consumers are classified and supported, or after an explicit documented scope decision. Keep release readiness distinct from local visual acceptance; the current issue remains open until required runtime/platform gates pass.

    P0 is the next implementation step. P1 and the P2 font-provider prototype follow its measurements. P3/P4 can provide savings before removing all Cairo metrics. P5 depends on proven target/dependency ownership; P6 depends on immutable state and a thread-safe metrics provider; P8 cannot pass until P1/P2/P7 eliminate mandatory native calls. Ship each step as a focused checkpoint with its own evidence and remaining limitations.

    Compatibility boundary for a literal 100% Cairo removal

    Managed mods compiled against Cairo.dll can keep working through a replacement compatibility assembly if its behavior is implemented. A mod that calls native Cairo directly, passes an external cairo_t*/surface pointer, or expects Context.Handle to be a real Cairo pointer cannot receive an opaque managed token and remain compatible. Options are to migrate that mod to a supported drawing API, implement a separately scoped native ABI adapter, or retain an optional native Cairo bridge. The bridge would defeat a universal no-Cairo claim for that mod. Do not silently reinterpret raw pointers or hide a native dependency in the fallback.

    There is no researched drop-in library that establishes complete compatibility for all 868 members, every platform surface, and external native pointer use. The concrete target is a Cairo-free client with a documented supported managed API surface. Full external native ABI emulation would be a separate large project and cannot be promised by adopting SkiaSharp or a text library.

    Measurement, correctness and finish criteria

    Use the normal authenticated datapath for runtime launches. Preserve credentials/settings and restore temporary harness changes. Use controlled fixtures and a disposable test world; record the exact runtime/source head, OS/driver/GPU, fonts, resolution, GUI scale, settings and workload. Never publish credentials or machine-specific paths. Keep screenshot/native-oracle work outside timed intervals.

    Measure cold open and warmed repeated use separately: unchanged dialog, ten matched full recompositions, rapid text edits, inventory hover/tooltips, scrolling/list menus, statbar updates, map/chat, rich text/item information, scale changes, SVG/tint updates and asset/font reload. Include gameplay with these events to expose CPU/GPU contention. Run at least five alternating A/B blocks after warmup; retain raw samples, medians, p95/p99, variance and sample sizes. Report main-thread CPU milliseconds per event, total process CPU work, blocked time, end-to-end latency, frame times, GPU duration, allocations/GC, upload/submission counts and cache behavior. A faster CPU thread with worse input latency is not sufficient. A busy GPU can also worsen world rendering.

    Proposed performance release targets, to be confirmed against P0 rather than treated as measurements: at least 30% lower warmed UI main-thread CPU time on the reference Linux workload, no material increase in total UI CPU work, and no more than 5% p95/p99 frame-time or end-to-end UI latency regression relative to the CPU baseline. Increase repetitions if noise crosses those limits. Do not report percentages for effectively zero-work cases; require zero unnecessary recomposition instead. Agree hardware-specific absolute budgets after baseline. Report failures rather than revising thresholds to accept the candidate.

    Correctness gates include all 57 override rows at scales 1.0/1.5/2.0, internal textures and direct SVG routes, unchanged shared visual policy, repeatable references, and negative comparator controls. Keep layout/interaction checks separate from antialiasing acceptance: a wrapping or hitbox error must fail even if pixels satisfy a tolerance. Extend text cases to all bundled fonts, fallback glyphs, combining marks, RTL/Arabic, CJK, whitespace, long words, emoji, mixed styles, selection and caret. Verify all implemented operators/clips/transforms, disposal/mutation/generation chains, queued writer closure, empty surfaces and large textures. Detect cold text atlas uploads and CPU glyph processing explicitly.

    For memory, track live targets, cached images/pictures/text runs, CPU bytes and driver-visible estimates separately. Stress at least 100 dialog open/edit/close and scale-change cycles, confirm references are released and memory plateaus after warmup, and repeat with a constrained memory budget. Handle context loss/reset by invalidating the entire context generation without stale texture reuse. Run GL sentinel/error recovery tests and preserve the existing ABI fixture. Clean bootstrap/build, patch/shader/compatibility gates and all five test projects remain required for implementation checkpoints. Windows/macOS validation is still pending for the latest implementation; library documentation is not a platform pass.

    The end state is GPU UI drawing/effects, retained resources, a Cairo-free metrics/state/software path, minimal measured CPU preparation and bounded asynchronous submission. CPU event logic, cache misses and driver command submission remain. A literal zero-CPU UI is not a completion claim.

    Optional research after the above gates

    If P0 still identifies a large parallel layout/geometry bottleneck after caching and removal of duplicate work, prototype a narrowly scoped GPU compute algorithm with explicit GPU-resident input/output and measured dispatch/readback costs. Evaluate dependency propagation, numerical precision and CPU hit-testing requirements. The current GL/SkiaSharp stack must expose the needed capability or use a separate compute API with validated synchronization. Stop if transfer/dispatch cost exceeds the saved work. This is optional research, not a dependency for Cairo removal or a promise that all layout can run through SkSL.

    For the next session, resume at feat/issue-130-retained-ui / 1383f3a, verify the remote head and any subsequent comments, and implement P0 counters and the matched baseline first. Attach each checkpoint's commit, commands, numerical results, native-call/module evidence, visual/input acceptance and unresolved items to this issue. The native Cairo oracle may remain in test tooling after the runtime dependency is removed.

  12. Zaldaryon commented on Oct 9, 2026

    @Zaldaryon
    CollaboratorAuthor

    Priority clarification: existing mod API compatibility comes first

    The compatibility requirement takes precedence over removing every native Cairo dependency. Existing modders should keep using their current Cairo commands and managed API without source changes, recompilation, a new drawing API, or mandatory GPU-specific setup. Preserve old compiled consumers as well as source compatibility. This refines the P0-P9 plan; it does not claim that complete parity has already been achieved.

    Keeping method names and signatures is only the first gate. Preserve assembly identity, public types, overloads, enum values, public fields, delegates and subclass/override points, plus observable behavior: text extents and advances, wrapping/autosize, font selection, current point, path queries, matrices and save/restore, source/pattern lifetime, clips, operators, antialiasing choices, surface formats/stride, pixel access and writes, exception/disposal behavior and thread ownership. Visual acceptance alone cannot approve a changed layout or hitbox. Cached or deferred rendering must not change callback timing, input behavior or the point at which a requested result becomes observable.

    Normal supported commands should execute through the replacement backend transparently. A command that cannot yet preserve its semantics should take a tested compatibility route rather than silently draw differently, return invented values, or throw merely because GPU mode is enabled. Explicit pixel access may require CPU materialization or GPU readback; count that cost and preserve the old contract. Do not represent an opaque managed token as a real Cairo native pointer.

    For direct native Cairo use or external native handles, retain an optional native compatibility bridge until equivalent support is proven. Managed commands should not require native Cairo once their replacement is complete. A process using the native bridge cannot be described as Cairo-free. If a rare API or platform surface cannot be preserved, document the exact exception and obtain an explicit scope decision before removing support. A forced mod migration is not the default resolution.

    Move the compatibility inventory and tests to the beginning of the plan, before replacing native state or metrics. Each API row should record the old contract, observed callers, backend implementation, fallback behavior and evidence. An API not seen in the sampled mods is not automatically safe to remove. Cover both ComposeElements and ComposeTextElements, custom contexts/surfaces, direct text and pattern calls, delegate drawing, and publicly mutable fields.

    Required per-step gates:

    • Run old compiled mod fixtures against the candidate, preserving the existing assembly ABI test and adding representative real consumers. Compile their source against the candidate too, to catch source-facing differences that an ABI manifest misses.
    • Compare return values, state transitions, lifetime, errors, output geometry, text metrics and requested pixels with native Cairo. Test mutation after binding, disposal with surviving references, save/restore nesting, direct handles, groups/masks, surface writes and font changes.
    • Exercise complete mod workflows on Linux first: opening dialogs, editing text, selection/caret, scrolling, tooltips, icons and repeated recomposition. Add Windows/macOS qualification before a cross-platform compatibility claim.
    • Keep unsupported operations covered by their compatibility route; verify GPU-off and initialization-failure modes. Fail a checkpoint on a changed old contract even when performance improves.

    The existing 81-type/868-member ABI inventory and old compiled fixture are starting evidence, not proof of complete behavioral compatibility. Acceptance requires every supported API row to have a verified implementation or a verified compatibility route. Preserve the existing shared visual criterion and add strict layout/input/state assertions where appropriate; do not widen tolerances to hide a mod regression.

    Implementation order is now: establish the mod compatibility contract and baseline, profile costs, replace internals behind that contract, then reduce the remaining native dependency wherever parity permits. The desired outcome is unchanged old mod commands with GPU rendering underneath. Complete native Cairo removal remains conditional on that compatibility requirement, rather than a reason to break existing mods.

  13. Zaldaryon commented on Oct 9, 2026

    @Zaldaryon
    CollaboratorAuthor

    Implemented the first, narrow P0 profiling slice in signed commit c9b9d77, on feat/issue-130-retained-ui.

    RecordingProfilingEnabled is independently opt-in and disabled by default. The backend now exposes cumulative counts/ticks for shadow creation and state-command replay, drawing capture, native path copying, native-to-Skia conversion and stroked-outline preparation. Reset preserves the enable flag; disabled scopes read no clock, update no counters and allocate no diagnostic object. Timings retain no UI content. The public harness README documents the API, scope boundaries and nested elapsed-time semantics. These are elapsed timings, not thread CPU or GPU measurements.

    Two focused tests preserve exact recorded pixels, current point and surface state with profiling on/off, validate operation counts and reset behavior, and verify zero counters when disabled. Fresh reconstruction from all public inputs applied 134/134 patches, built with zero warnings/errors, passed all five projects with 1,141 passed and 34 existing skips, and passed patch/shader/compatibility gates. Compatibility retains 17 documented skips. The existing old compiled Cairo mod fixture also passed. No drawing API or text/layout behavior was replaced.

    Pharos passed six completed inventory scenarios on Linux with the physical GPU and standard authenticated datapath. Profiling is enabled only around the measured recompositions and restored afterward; existing JSON settings were restored after every run. The controlled workload uses two warmups, then ten inventory recompositions. One complete comparison block produced:

    Mode Elapsed time for ten recompositions
    Published GPU baseline, before instrumentation 878.02 ms
    Candidate GPU, profiling disabled 951.58 ms
    Candidate GPU, profiling enabled 1,005.29 ms
    Candidate CPU renderer 491.72 ms

    One additional CPU run measured 530.20 ms. The broader five-block campaign was stopped between launches at the request to finalize; these samples do not establish profiling overhead, CPU savings or a speed improvement. The higher disabled-candidate time is an unresolved measurement signal, not evidence that instrumentation is free. Repeated alternating controls are still required.

    The controlled profiled GPU block recorded 350 shadow creations, 2,910 shadow accesses, 13,870 completed replay commands, 2,160 drawing captures, 1,720 path copies/conversions and 980 stroke outlines:

    Preparation stage Inclusive elapsed duration
    Shadow creation and replay 31.963 ms
    Drawing capture 464.359 ms
    Native path copy plus conversion 134.706 ms
    Native-to-Skia conversion alone 10.439 ms
    Stroke-outline preparation 3.721 ms

    Drawing capture includes most path copying and outline preparation; path copying includes conversion. Do not sum the nested rows. Clips can copy paths outside drawing capture. These measurements place capture well ahead of shadow replay in this workload, so the next profiling step should attribute capture's remaining state/source/native-copy work before selecting an optimization. Text measurement and text-to-path construction before capture are outside this slice. Actual main-thread CPU attribution, GPU timing and the full P0 baseline remain open.

    All measured GPU intervals retained zero CPU rasterization, native materialization, diagnostic readbacks and fallback. Existing scale fixtures and GL recovery checks passed. This checkpoint adds evidence for the compatibility-first plan; it does not establish complete behavioral parity or remove native Cairo. Windows/macOS checks remain pending for this change. No PR or merge is included; the issue remains open.

  14. Zaldaryon commented on Oct 9, 2026

    @Zaldaryon
    CollaboratorAuthor

    Implemented immutable preparation reuse on feat/issue-130-retained-ui in 1912952.

    GPU recording now reuses text/glyph outlines and advances, native-to-Skia path conversion, stroked outlines and NanoSVG pictures. Keys compare the complete inputs after hashing. Text entries retain their scaled font, paths are copied before mutation, and recorded SVG streams retain their own resource leases. The shared LRU is capped at 1,024 entries and a 16 MiB estimated preparation-resource budget, with a 128 KiB key limit. Native font internals can consume additional memory, so that estimate is not a process-memory ceiling.

    Compatibility remains the boundary. BeforeCalcBounds, ComposeElements, ComposeTextElements and OnComposed still run. Each draw reads fresh paint, source, alpha, clip and transform state. Empty or malformed text, unsuccessful contexts, user fonts and oversized inputs use the native preparation route. Native command replay remains the visual oracle. The implementation preserves the old mod commands and does not remove Cairo.

    Validation on Linux with Vintage Story 1.22.7 and Pharos v0.5.0:

    • Clean public-input reconstruction, build and IL patching passed with zero compiler warnings/errors. All five test projects passed: 1,168 passed, 34 existing skips, zero failures. The change adds 27 preparation tests, including input mutations, path/current-point continuity, glyphs, explicit font matrices, malformed input, native error state, hash collisions, entry/byte eviction, live resource ownership and concurrent producers.
    • Patch checks passed for 81 source patches and 53 Cecil patches, with zero pending, unavailable or conflicting patches and 22 runtime donors compiled. Shader and multiplayer compatibility checks passed, with 17 existing allowlisted skips and zero cast divergences.
    • The Cairo public/protected inventory remains identical at 81 types and 868 members. A mod fixture compiled against the pre-existing Cairo binary ran against the candidate assembly.
    • Final candidate binaries passed all three Pharos functional scenarios. The changed-input fixture passed 36 scale/mutation cases with byte-identical GPU pixels for cache off, cold and warm runs, plus all old custom-draw and OnComposed callback counts. Its 228 GPU/native prefix comparisons passed. The full existing matrix passed 171/171 override/scale cases across 1,955 captures, under the unchanged linux-gpu-cairo-local-aa-v2 policy and all comparator counterexamples. Native reference repeats were identical. Exact runtime and test-host binary hashes match the clean candidate.
    • Separate profiled controls recorded path conversion counts of 1,720 to zero and stroke-outline preparation counts of 980 to zero after warmup. Native path copies still occurred (1,720 to 1,400). Capture measured 433.06 to 412.61 ms in those diagnostic runs; those timings are not the unprofiled performance result.
    • Five matched unprofiled GPU off/on pairs used two warmups and ten inventory recompositions per process, with the pair order alternated. Median wall time changed from 962.15 to 946.74 ms (-1.6%); median render-thread CPU time changed from 926.98 to 916.33 ms (-1.1%); render-thread managed allocations changed from 399.14 to 379.99 MB (-4.8%). The cache-on means are different: wall time increased 2.7%, while thread CPU time fell 2.9%. One cache-on wall-time outlier of 1,284.58 ms is included. The paired mean CPU saving is 26.8 ms, with a 95% t interval of -42.9 to +96.4 ms, so these five pairs do not establish a reliable CPU-speed improvement or reduced stutter.
    • Every warm cache-on measurement recorded 2,700 hits (320 text, 1,400 path and 980 stroke), zero misses/evictions, 194 entries and 1,396,557 estimated bytes. Normal measured composition recorded zero CPU rasterization, native materialization and diagnostic readback, with no GPU fallback. These counts establish preparation reuse independently of timing variation.
    • Two unchanged published-GPU controls at c9b9d77 measured medians of 956.82 ms wall / 926.29 ms thread CPU. Two native-Cairo CPU controls measured 482.91 / 465.44 ms and 187.97 MB allocations. GPU recording still costs about twice the render-thread CPU time of native Cairo in this workload; this checkpoint does not resolve that regression.

    The runtime exposes RecordingPreparationReuseEnabled, cache clearing/reset methods and hit/miss/eviction counters for matched controls. The harness README documents the bounds, ownership and measurement procedure. Profiling remains a separate opt-in control; preparation counters now reflect actual misses.

    The next optimization should profile the remaining paint/source capture, native shadow/path copying, text measurement/layout and GPU blur/submission costs. Exact text-metric reuse and source preparation need separate mutation and ownership tests. Layout and arbitrary mod callbacks still execute on the CPU. Skipping a whole element without an explicit invalidation contract can miss external mod state or callback effects. Removing Cairo remains conditional on preserving both the old binary API and these behaviors. This checkpoint does not establish universal mod compatibility or a gameplay frame-time improvement. Issue #130 remains open.

  15. Zaldaryon commented on Oct 10, 2026

    @Zaldaryon
    CollaboratorAuthor

    Published the Windows profiling and allocation-tracing checkpoint in signed commit 6b4f4e4, on feat/issue-130-retained-ui.

    The remaining capture cost was dominated by Cairo allocation stack traces. The fork forced CAIRO_DEBUG_DISPOSE=0 and enabled tracing whenever the variable existed. The correction preserves the caller's environment, disables tracing for unset/0, and keeps the public mutable CairoDebug.Enabled field usable after initialization. Native drawing, text metrics, mutable source state and mod callbacks retain their existing routes. Added opt-in stage diagnostics cover state/source capture, text/font work, blur, replay and asynchronous submission; nested host elapsed times must not be summed or presented as completed GPU execution.

    Windows measurements used Vintage Story 1.22.7, Pharos 0.5.0 and a physical Radeon RX 6650 XT. Each control ran seven alternating Cairo/GPU pairs with preparation reuse enabled, three warmups and ten measured inventory recompositions at scale 1. Profiling and composition tracing were disabled. Settings and native library hashes matched; both runs left the debug environment and field override unset. Harness assemblies were rebuilt, so these are equivalent workloads rather than identical harness binaries.

    Median per ten recompositions Original GPU Candidate GPU Candidate Cairo
    Render-thread CPU 625 ms 312.5 ms 359.375 ms
    Wall time 643.09 ms 320.49 ms 358.91 ms
    Render-thread managed allocations 289.09 MB 112.38 MB 139.13 MB

    Observed GPU CPU fell 50% and managed allocations fell about 61% in this warmed inventory workload. This does not establish a universal speedup or gameplay frame-time improvement. All seven candidate GPU intervals recorded zero fallback, native materialization and diagnostic readback. Separate diagnostic pairs measured about 5.8 ms source capture, 7.2 ms path copying, 12.8 ms native text/font measurement, 26.8 ms blur host work, 39.8 ms replay host work and 31.0 ms flush/submission. OpenGL elapsed-query results varied substantially on this driver and do not support completed GPU-time claims.

    Windows validation:

    • Fresh reconstruction applied 135 patches with zero skipped/failed. Release build passed with zero errors and 101 existing warnings. All five test projects passed: 1,189 passed, 34 explicit existing skips, zero failures.
    • Strict patch checks passed: 82 source patches, 53 Cecil-owned patches, zero pending/unavailable/conflicts, plus 22 runtime patches applied and exact donors compiled. Shader checks passed; multiplayer compatibility checks passed with 17 documented skips.
    • Cairo public/protected ABI remains identical at 81 types and 868 members. The old compiled Cairo fixture ran against the candidate. All 33 bundled-font scale fixtures retained native metrics and recorder fallback pixels.
    • Pharos passed the real-client GUI matrix: 66 captures across 11 states and three scales, no element-bound mismatch, client error or tracing gap. This exercises 26/57 manifest rows at scale 1 and 32/57 at scales 1.5/2, not the complete Linux visual matrix. Separate GL checks passed 4,096 candidate/state/fallback cycles and 32 context recreations.

    Windows GPU compatibility is still incomplete. The strict partial-blur fixture fails identically on the original checkpoint and this candidate: 32x24, range 0, edge 4, translucent input, maximum RGB difference 28 and alpha difference 0. The native algorithm intentionally skips accumulator updates and wraps byte values, so range 0 is not an identity operation. Tolerance was not relaxed. The next step is a Pharos reproduction with intermediate horizontal/vertical byte comparisons, followed by the smallest proven compatibility fix.

    macOS: skipped, no hardware. Linux was not rerun in this Windows session. Issue #130 remains open; this checkpoint does not remove Cairo or establish complete mod compatibility.

  16. Zaldaryon commented on Oct 10, 2026

    @Zaldaryon
    CollaboratorAuthor

    Follow-up pushed in 619d17d. The strict Windows partial-blur failure is fixed. The recorded source rounded straight RGBA to bytes before premultiplication; its one-level input differences became a 28-level RGB difference under native partial-blur byte wrapping. Solid paints now preserve Cairo's premultiplied 16-bit-to-8-bit conversion. The blur shader and acceptance thresholds are unchanged.

    The operator controls also exposed a preexisting coverage defect: masked Source could lose 127 alpha levels over an opaque background. Masked Clear, Source, In, Out, DestIn and DestAtop now use native replay, following Cairo's classification of operators not bounded by source. This preserves the uncovered destination and intentionally gives up GPU eligibility for those cases.

    Windows validation ran as Pharos ClientServerScenario tests on the real client render thread with the Radeon RX 6650 XT. All 120 full/partial blur cases passed with 12 GPU shader passes each; observed maximum RGB and alpha differences were both zero. The 56 operator/background/mask scenarios included 36 exact native-fallback checks. The same client passed 4,096 candidate, wrong-thread, injected-error, GL-state and texture-deletion cycles, plus 32 Skia context reinitializations on its existing GL context. The client continued rendering afterward with no logged errors. The GUI/interaction matrix passed with 66 captures across 11 states, three scales and both render modes.

    Fresh Windows reconstruction applied all 135 patches without failure. Build, strict patch validation, runtime donor compilation, shader and vanilla compatibility checks passed. The five test projects report 1,197 passed, 34 existing explicit skips and zero failures. The old compiled Cairo fixture still runs, and all 81 public types and 868 public/protected members retain their original identity and signatures.

    The final unprofiled inventory control used seven alternating Cairo/GPU pairs, three warmups and ten measured recompositions. Median render-thread CPU was 296.875 ms with GPU and 343.75 ms with Cairo; managed allocations were 112,372,792 and 139,129,040 bytes respectively. This is one measured inventory workload, not a general gameplay speed claim.

    The issue remains open for the remaining primitive and mod coverage. Masked Over retains the previous output but still differs from native by one or two channel levels; gradients remain separate diagnostics. Invoked GUI manifest coverage was 26/57 rows at scale 1 and 32/57 at scales 1.5 and 2. The 1280x720 viewport clips oversized dialogs in both modes. macOS: skipped, no hardware. Linux was not rerun in this Windows session.

  17. Zaldaryon commented on Oct 10, 2026

    @Zaldaryon
    CollaboratorAuthor

    Pushed 7e0da04 on feat/issue-130-retained-ui, following 619d17d: masked solid colors now clamp opacity before multiplying source alpha and use Cairo's premultiplied quantization. The real Pharos native/GPU comparison caught mask 1.25 changing RGB by up to 49 levels and alpha by 51. All 18 tested mask/backdrop cases now preserve exact source bytes on transparent pixels and stay within one channel level on nontransparent backdrops.

    The independent 96-case gradient matrix exposed large coverage differences for None and Repeat, including native endpoint colors becoming transparent or wrapping to the first stop. Linear and radial gradients in those two modes now use the existing native replay fallback. Its 48 cases match native bytes exactly. Pad and Reflect remain accelerated; their 48 GPU cases stay within one level per channel. Mutation and disposal after drawing are included. Native fallback cases are not GPU parity passes.

    A separate mod-shaped dialog passed real mouse/keyboard number and multiline text input, static/dynamic drawing callbacks, Compose/ReCompose/Redraw, BeforeCalcBounds/OnComposed, custom render callbacks, image/background rendering and identical bounds at scales 1, 1.5 and 2 with GPU off/on. Each GPU composition variant recorded 19 successful GPU textures with no materialization or fallback. This adds bounded functional coverage for five previously missing element types, without claiming acceptance of every GUI fixture or third-party mod.

    Validation on Windows 1.22.7 with unchanged Pharos 0.5:

    • Clean bootstrap: 135 patches applied, no skips or failures. Full solution: zero errors, 101 existing warnings.
    • Five unit suites: 1,207 passed, 34 existing skips, zero failures. Strict source/Cecil/runtime patch checks, shaders and vanilla compatibility passed; the compatibility check reports 17 explicit skips.
    • Real Pharos: four compatibility scenarios plus a separate inventory control passed. The existing 120 blur cases, 4,096 GL state/error/thread/deletion cycles and 32 Skia context reinitializations on the same live GL context remain healthy. Original GUI matrix: 66 captures; mod dialog: six more. No client errors, bounds mismatches or trace gaps. Oversized dialogs still clip in the 1280x720 viewport at larger scales in both modes.
    • Old compiled Cairo consumer ran unchanged. Public inventory remains identical at 81 types and 868 members.

    Seven alternating inventory pairs, three warmups and ten measured recompositions, scale 1, profiling/disposal tracing disabled, with no concurrent build or patch extraction: GPU median CPU 328.125 ms, wall 322.3567 ms and 112,361,856 allocated bytes; Cairo CPU 359.375 ms, wall 357.8872 ms and 139,130,160 bytes. This workload reported no materialization, readback or fallback. These are local medians, not a universal speed claim; the cost of the newly native gradient modes is not established by this inventory scenario.

    macOS: skipped, no hardware. Linux was not rerun in this Windows session. NaN opacity remains unverified. Next: compare NanoSVG gradient spread/opacity against its own native rasterizer in Pharos, then continue the remaining GUI fixture coverage. NanoSVG uses a separate source contract, so the Cairo gradient fallback was not copied into that path.

  18. Zaldaryon commented on Oct 10, 2026

    @Zaldaryon
    CollaboratorAuthor

    Pushed 087bf5f to feat/issue-130-retained-ui.

    The NanoSVG follow-up fixes solid paint opacity/premultiplication and reconstructs integer color bytes before SVG overlay composition. Ordinary image rounding stays unchanged. SVG recording is limited to at most 256 disjoint pixel-aligned rectangles with complete texture coverage. Gradients, curves, strokes, overlaps, uncovered margins and partial zero-alpha tiles use the bundled native renderer. This gives up acceleration for those SVGs to preserve their pixels; it does not establish compatibility for every mod.

    Real Pharos on Windows 1.22.7/Radeon RX 6650 XT passed 3,240 SVG comparisons: 360 actual GPU cases within one RGB byte and exact alpha, plus 2,880 explicit native fallbacks with exact bytes. The matrix includes tint, near-zero opacity, three GUI scales, gradient spread/focus/transforms and representative icon geometry. The four surrounding Pharos scenarios also passed: source compatibility, mod callbacks/input, GUI/recomposition, and strict blur/OpenGL recovery. A decimal-formatting error in the expanded SVG fixture was corrected and that scenario passed separately. The surrounding run includes 120 strict blur cases, 4,096 GL state/error/deletion cycles, 32 Skia context reinitializations and 72 GUI captures.

    Clean bootstrap applied all 135 patches. Five unit suites passed 1,226 tests with 34 existing skips; strict source/Cecil/runtime patch, shader, vanilla compatibility and legacy ABI checks passed. The public Cairo inventory remains 81 types and 868 members. macOS: skipped, no hardware. Linux was not rerun in this Windows session.

    The isolated inventory control used seven alternating pairs, three warmups and ten recompositions at scale 1 with profiling/debug disabled. Median GPU CPU/wall time was 375/487.8 ms versus Cairo 390.6/449.6 ms; allocations were 112.36/139.13 MB, with zero materializations, readbacks or fallbacks. CPU timing is coarse and wall timing remains variable, so this does not demonstrate a general speedup. The next performance investigation is the remaining host/submission wait and the preparation cost of native SVG fallbacks.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions