THE ZLANG OBSERVATORY
A language taking shape.
Streaming by nature. Responsive by design.
An independent language for fast applications and first-class zui.
Loading checkpoint…
Published snapshots · refreshes every 60sOn the workbench
IN PROGRESSLoading current work…
Browse checkpoint ↗PERFORMANCE LAB
Less work. Measured honestly.
Explore the cost of state updates, then compare generated code with Rust. Every result keeps its scope and source checkpoint.
One service. Different subscriptions.
How broad and granular invalidation trade CPU for memory.
Loading workload…
Update CPU
ns / frameMedian ± median absolute deviation
Retained heap
KiB · requested live Rust allocation bytes
Rust callbacks and native state tracking. Excludes GPUI rendering, Text, I/O and Zeron. This compares policies, not zlang against Rust.
All 8 workloads & measurement details
Making bursts cheaper
Before & after · skip subscriber walks when every view is already dirty
MEASURED CHANGE
No extra tracking metadata. One clean unrelated view disables this global shortcut; it does not solve general subtree rendering.
Refresh the mounts that changed
Two measured steps · coalesced wakes, then fewer child synchronization passes
Refresh CPU
µs / frameMedian ± MAD · not a confidence interval
Retained setup heap
KiBRequested live Rust bytes · not RSS
Actual mounted-tree adapter and native bindings. Mixed build profile: adapter/native/runtime optimized; other dependencies use the dev profile. Excludes GPUI layout, paint and whole-application performance. Comparisons were timed separately; do not multiply stage ratios.
All 31 workloads & source checkpoints
The cost of subtree caching
Experimental component rendering · direct renderer stays the default
Frame CPU
ms / requested frameMemory cost
Releasing removed layout contexts
Paired private dependency fix · separate from cache policy
Frame CPU
ms / requested frameMedian ± MAD · not a confidence interval
Retained requested heap
KiB after updatesThe fix drops contexts for removed layout nodes. Cached rendering still retains more heap than direct rendering. This is successful Rust allocator requests, not RSS or a whole-process memory limit; CPU results exclude GPU submission and rasterization. Neither cache promotion nor Rust parity follows from this ownership fix.
One record, independent field reads
Native Int cells · coarse revision, independent cells and a record cohort
Update CPU
ns / logical frameMedian ± MAD · not a confidence interval
Requested heap
Sparse subscriptions avoid unrelated evaluations; dense updates can cost more than coarse invalidation. Record aliases retain the whole record, while independent cell aliases have independent lifetimes. This measures Rust callbacks over flat Int fields, not generated components, Text, GPUI rendering or language parity. Dense writes update fields individually, not as an atomic batch.
Safer wake delivery, measured costs
Native state notifications · two separate paired studies
Native CPU
ns / iterationMedian ± MAD · not a confidence interval
Native memory
Each study uses its own paired baseline. Historical timings are never pooled across studies. Native callback costs exclude generated handlers, async scheduling, GPUI layout/paint, and whole-application Rust parity.
The cost of safe row lifetimes
Native key windows and removal validity · measured tradeoffs
Operation CPU
ns / operationMedian ± MAD · lower is better
Requested heap
Seven randomized paired rounds on one affinity CPU; setup and teardown are outside timing except registration churn. Allocation runs are separate. Registrations are validity guards, not expression watchers or rendered row scopes. Unchanged controls also vary; this is not a CPU win or service-driven UI benchmark. No generated repetition, GPUI layout/paint, GPU or Rust parity claim.
Measurement scope and retained adverse result
The initial inline map retained 368 bytes after its final registration. That original report remains in the private repository. Both final variants release empty map storage; lazy storage adds one 24-byte box while registered. More frequent release increases churn allocation traffic.
Two viewports, three rendering policies
Generated service rows · CPU and requested heap tradeoffs
CPU per iteration
Median ± MAD, not a confidence interval. Above 1× direct is slower.
Requested Rust heap
Instrumented stage calls across all measured iterations · left / right. Layout, prepaint and paint count pane boundaries, not every descendant. Scroll horizontally for all stages on small screens.
| Policy | Renders | Rows | Layout | Prepaint | Paint |
|---|
All three policies use the SAME fine-grained native dependency machinery. This compares renderer, cache and invalidation choices, not alternative native graphs. Actual manually driven headless GPUI layout, prepaint and paint are included; OS wake scheduling, GPU work and the surrounding parent visual tree are excluded. Requested Rust heap is instrumented separately from timing; it excludes RSS, C allocations and GPU memory. No equivalent Rust application baseline or language parity claim.
Frozen source and measurement scope
Snapshot JSON: raw samples, counters and source hashes ↗COMPLETE GENERATED PARENT · SEPTEMBER 2026
Keep the clean pane. Measure what stays.
Inline, sibling caches and sibling + pane cache · one shared provided List
Keeping the clean pane reduces total CPU for larger Footer updates in this fixture. Small payload updates regress, and lower allocation traffic comes with higher retained heap.
Purple: Inline · Mint: Siblings = ordinary sibling caches · Amber: + Pane = sibling caches plus clean-pane caching
Update CPU
µs / update32 warmup updates, then 256 measured updates per process · 7 timing processes per policy · median ± unscaled MAD
Paired CPU ratios
above 1× is slowerPair policies within each round, then summarize. Ratios need not equal ratios of independent CPU medians; MAD is not a confidence interval.
Interval allocation traffic
One separate allocator process per policy · interval totals ÷ 256 · excludes intervening verification · no dispersion estimate
Instrumented live heap
This is the same complete generated parent and one shared List, with actual 1 or 6 mounted rows. No adaptive row cache, whole-service broad invalidation or Rust application comparator. Manual headless GPUI layout/prepaint/paint excludes GPU presentation and automatic X11 wake latency. Separately retained X11 correctness cases are not these CPU measurements. No production default is promoted.
Exact CPU, allocation, heap and measurement method
CPU and ratio cells show unrounded median ± MAD. A dash means a zero comparator denominator. Allocation cells are exact interval totals ÷ 256 from one instrumented process.
FOLLOW-UP · NORMAL APP EFFECT BOUNDARY
One less creation. No consistent CPU gain.
Fresh sibling subtree reuse · baseline versus first-prepaint candidate
The candidate saves a small amount of Footer allocation traffic. Paired CPU results do not show a consistent improvement; the 4,096-row payload case with pane caching is slower. Unchanged inline controls stay visible to show run variation.
Candidate / baseline CPU
above 1× is slowerAll policies shown. Inline is unchanged by the optimization. Median of seven adjacent within-round ratios ± unscaled MAD; not a ratio of medians or a confidence interval.
CPU per update
µs · median ± MADPurple: baseline · Mint: candidate. Six clock endpoints remain included. The boundary residual includes App effects, callback boundary and extra clock overhead; it is not pure effect-flush time.
Allocation traffic
Separate instrumented twins: one process per variant/context, no dispersion estimate. Interval excludes normal App effects; effect-inclusive includes them. Both traffic intervals exclude intervening verification.
Normalized requested heap
Each update returns through the normal outer App effect boundary. This differs from the older study's single enclosing callback.
Process-global requested Rust heap; includes outside-timing verification. Before-model is after platform setup. Not RSS, GPU memory or timing-build heap. Equal warm/end endpoints do not prove forever-bounded memory or a production memory win.
32 warmups, then 256 measured updates; CPU 0, seven paired rounds. Timing compiles allocator/lifecycle/cache/stage probes out, but retains one uniform per-Window draw-entry/call counter. The legacy “instrumented = false” flag refers to allocator instrumentation only. A draw-entry count is not independent proof of completed paint. Native values, identities and mounted row bounds are checked outside timing; allocator twins also check full-parent geometry.
Total surrounds writes, routing, explicit GPUI draw and the normal App effect flush. This manually drawn headless fixture uses one complete generated parent and one shared List; it does not measure automatic platform frame scheduling, GPU presentation, adaptive caching, whole-service invalidation or a Rust comparator. Known builds/tests were held during collection; unrelated shared-host activity was uncontrolled. No CPU win is inferred from allocation savings.
Exact CPU phases, paired ratios, traffic and normalized heap
MATCHED ROW STRATEGIES · SEPTEMBER 2026
Faster sparse rows. A cost to rebuilding.
Raw rows, corrected nonadaptive cache and adaptive v3 · identical outer panes
Adaptive v3 reduces sparse drawing work against raw rows in this fixture. It does not consistently beat corrected caching: cold priming, repeated eviction and small panes expose the tradeoffs.
Purple: Raw = raw row elements · Mint: Fixed = corrected nonadaptive cache · Amber: V3 = adaptive v3
Absolute CPU
µs / phase sampleIndependent medians ± unscaled MAD · 7 processes per strategy
Paired CPU ratios
above 1× is slowerPair within each round, then summarize. These ratios need not equal the ratios of the absolute bars.
Allocation traffic
One separate instrumented process per strategy · phase average, no dispersion estimate
Requested live heap
All three strategies share cached-granular outer panes and corrected-abort GPUI/native inputs. Raw rows are not whole-service invalidation or a Rust comparator. Explicit headless draws measure layout, prepaint and paint CPU; no GPU submission/rasterization or real X11 window loop is measured. No strategy is promoted.
Exact selected values, all phases and measurement method
CPU cells show unrounded median ± MAD in nanoseconds per phase sample; ratios show unrounded median ± MAD. A dash means the denominator is zero. Allocation cells are exact per-phase averages from one process, not medians. Heap snapshots are bytes normalized to each process before app/model setup.
What cache eviction costs
Follow cold starts, repeated changes and the frames after eviction
Cost through successive frames
● Corrected cache, retained ● Adaptive candidate · hover a point for its exact value
Each sequence is one process per variant, not a timing distribution. These are actual headless GPUI frame allocations. Version 2 requires two clean readers and half the pane clean; version 3 retains caching with any clean reader. Both defer eviction until the third successfully published ineligible frame. Cold priming and the cost of rebuilding after eviction remain visible. Neither candidate is enabled.
Live memory subtracts each process baseline after model creation, before opening the window. It includes GPUI/font caches and excludes allocator metadata, RSS and GPU memory. The baseline is the corrected nonadaptive row cache, not the original direct renderer. Configured viewport counts are checked against actual mounted rows for the plotted non-cold phases. No automatic window scheduling, full-parent cost or Rust parity is measured here.
Exact phase values and provenance
Reusing clean visible rows
Sparse wins, churn costs · measured CPU and retained memory
These timings measure the earlier cache: that version has an ancestor prepaint rollback defect. A separate corrected candidate passes retry tests, but its extra token and observer costs are not measured here. This cache remains experimental; corrected and adaptive performance acceptance is open.
Frame update CPU
µs / iterationMedian ± MAD · seven paired rounds
Memory cost
51 matched service cases preserve measured row geometry, values and native cleanup. Sparse 4096-row granular updates use about 0.325× CPU with 19,900 extra warm retained bytes. All-visible churn regresses to 1.057× CPU; the one-row sparse case shows no established win. Larger panes mount eight visible rows each. Dense changes the first row in each pane; all-visible changes all eight.
Headless GPUI layout/prepaint/scene work; excludes GPU rasterization and OS scheduling. CPU and requested Rust heap use separate binaries. Nine missing font families precede the fallback in this environment. This is neither an enabled optimization nor a Rust-parity result.
Study provenance
Where frame memory goes
Smaller temporary layout arrays · same service workload
Temporary Taffy layout arrays account for 84.6% of requested bytes in the large sparse granular case. Reserving smaller arrays reduces those bytes; it does not remove allocation calls.
Frame update CPU
µs / iterationMedian ± MAD · seven paired rounds
Memory cost
All 42 service cases preserve geometry, values and native cleanup; 840 differential layouts agree. CPU results are mixed, including small-view regressions and noisy unchanged controls. No retained-memory win is demonstrated. Dense updates change the first row in both panes, not every visible row.
Headless GPUI layout/prepaint/scene work; excludes GPU rasterization and OS scheduling. CPU and requested Rust heap use separate binaries. Nine missing font families precede the fallback in this environment. This is neither an enabled optimization nor a Rust-parity result.
Study provenance
What native collections actually cost
Production ABI · checked owners, stable rows and reactive identities
Operation CPU
ns / operationMedian ± MAD · lower is better
Requested heap
Five randomized paired rounds on one pinned CPU; setup, checksum and teardown are outside the timer. Allocation runs are separate. Indexed native reads include checked lookup plus row read; Vec/Deque use direct indexing and omit stable handles, shared ownership and subscriptions. Alias cases compare retain/release with Rc<Vec> clone/drop only. Native-only observer and Text lifetime probes remain in the raw snapshot. No compiler, async, GPUI or whole-application parity claim.
Finding the right collection storage
Paired storage prototypes · adverse results remain visible
Operation CPU
ns / operationMedian ± MAD · lower is better
Requested heap
Vec and VecDeque are lower-level Rust references: they preserve inline values and insertion identities, but omit shared ownership, reactive subscriptions and runtime schema/stale-handle checks. This is not a complete runtime parity comparison. The original prototype allocates temporary row values during append; the compact prototype uses borrowed scalar buffers. Wide middle moves benefit from moving keys instead of entire records. Text and subscription lifetime probes remain in the raw report.
zlang × Rust
Checkpoint microbenchmarks · time & allocations
Overall Rust parity remains open. Small UI differences are within measurement noise; these results do not measure a complete application.
Source provenance & limitations
THE ROAD AHEAD
A clear view of the work.
Tracked tasks, not a language completion percentage. The list grows as implementation and testing expose more work.
Language foundations
EVIDENCE, NOT ASSUMPTIONS
Built. Tested. Still being verified.
Passing tests cover the implemented subset. Whole-language Rust-equivalent safety and application performance remain unverified.
CHECKPOINT LOG