Local Linux HTTP comparison, October 1, 2026

CN1 source: 60c4f310bddcfc51e5acc75c63a57c43c73430d3
Apple M4 Max host (16 logical CPUs, 64 GiB), Apple Virtualization Linux ARM64 VM
with 4 vCPUs and 8 GiB. Fedora 41 host kernel, Debian 12 benchmark container.
All native compilation and measurements are LOCAL. No native CI was used.

Applications:
- CN1 lower-level Bench handler, -O3 glibc ARM64 backend package settings.
  BENCH_REUSE_RESPONSE=0 allocates each response; BENCH_JSON_MODE=0 builds a fresh
  LinkedHashMap for each JSON response. No annotation or transaction work in this
  handler. Default packaged 4 MiB collector trigger; native virtual-thread path.
- Spring Boot 4.1.1, spring-boot-starter-webmvc, embedded Tomcat defaults.
  JDK: Temurin 25.0.4.1+1. No JVM AOT cache, explicit GC choice or heap cap.
- The same Spring application: Oracle GraalVM 25.0.4+7.1 Native Image, default
  Maven native profile, no PGO or march=native. Default optimization level.

This compares the selected HTTP stacks, not equal frameworks or annotation
processing. Plaintext is 13 bytes. JSON has a message field with Hello, World!.
CN1 uses a LinkedHashMap and Spring uses Map.of. No TLS, database, authentication,
telemetry, sessions or realistic business logic. See sources, flags and raw logs.

Method fixed before the matrix:
3 rotating/interleaved rounds, 3 runtimes, 1/2/4 allowed server CPUs = 27 fresh
processes. Round 1 reverses the CPU order; each round rotates runtime order.
taskset confines all server threads to CPUs 0, 0-1 or 0-3. CN1 WORKERS matches
that count. At 1/2 CPUs, wrk runs on remaining VM CPUs; at 4 it shares them.
The 4-core points are contention experiments, not clean server-scaling results.
All services share the same 8 GiB VM; there is no per-service memory cap.

Startup: process spawn to first validated HTTP response. Polling sleeps 1 ms
between failed attempts. Filesystem caches are not flushed. Includes taskset
and readiness-probe overhead; not a cold-machine or serverless measurement.

RSS: /proc/PID/status. Idle before load is measured 3 seconds after readiness.
Peak is VmHWM after the two warmups and measurement windows, before the final
5-second idle sample. Measurement-window RSS is sampled every 50 ms too.
Peak includes startup, not just the traffic interval. None of these is live heap.

Each route: validate status and body, 10-second wrk warmup, 10-second measurement,
2 load threads, 32 connections, plaintext then JSON. Both keepalive limits are
raised/disabled (CN1 burst 1000000000, Tomcat max-keep-alive-requests=-1).
Any socket/non-2xx report in warmup or measured output aborts the run.
All 27 accepted processes passed these checks. Short warmup does NOT establish
steady-state JVM throughput; differences between routes can include warmup.

Summary statistics are medians across the three fresh processes per cell.
Throughput graph whiskers show min/max. Complete per-run results are preserved.
The featured 2-CPU configuration reserves other VM CPUs for the load generator.
Results are exploratory local measurements, not production capacity estimates.

Size: actual executable bytes from default native build outputs; no compression.
CN1's standard linker strips symbols. Shared libraries are excluded, and the
linked dependency lists are in toolchains.txt. Not container/deployment totals.
JVM native executable is N/A; the Spring JAR is not a standalone native binary.
The stripped-sizes file, when present, gives an additional strip-both comparison.

Reproduce in an isolated Linux ARM64 VM with 4 vCPUs and 8 GiB:

1. Check out the CN1 commit above. Set JDK_8_HOME, JAVA_HOME and PATH to Java 8.
   CN1_BACKEND_STANDALONE_DEMO=1 CN1_BACKEND_DEMO=demo/bench \
     ./vm/backend/package.sh Bench com.demo glibc-arm64
   This translates with Java 8 and links through the local container engine.
   Copy vm/backend/target/dist/bench-linux-glibc-arm64 into /out/cn1-bench in your
   benchmark container based on cn1-backend-glibc-arm64. Create /out first.

2. Install gcc, maven, python3, wrk, procps and util-linux inside that container.
   Put Linux ARM64 Temurin JDK 25 at /opt/jdk25 and GraalVM 25 at /opt/java.
   Put App.java in /spring/src/main/java/example/App.java and spring-pom.xml in
   /spring/pom.xml. Build with the matching GraalVM JDK (the application requires
   Java 25; the CN1 repository build in step 1 uses Java 8):
   cd /spring
   JAVA_HOME=/opt/java PATH=/opt/java/bin:$PATH mvn -Pnative package native:compile

3. Copy linux/benchmark_linux.py into the container and run it with Python 3.
   It writes /out/results. CPU IDs 0-3 must be available. Stop unrelated workloads
   before measuring; do not run the applications concurrently.

GraalVM Linux toolchain download used:
https://download.oracle.com/graalvm/25/latest/graalvm-jdk-25_linux-aarch64_bin.tar.gz
Pin versions from toolchains.txt when reproducing later. Artifact hashes are in
linux/results/results.json. With Python, matplotlib and numpy installed, run
plot_benchmarks.py and plot_final_metrics.py from any directory to recreate the
HTTP charts. Outputs go to charts/native-http inside the extracted bundle.

The superseded-macos-experiment folder retains the preliminary platform-thread
experiment, including its different response mode and short warmup. Its results
are not mixed into the Linux tables. The initial default-burst Mac trial failed
with read errors and was rejected; the older README records that limitation.

transaction-example contains the separate controller and service source shown in
the article, plus illustrative transaction pseudocode. The benchmark does not
exercise that service. The pseudocode explains the transaction boundary, not a
generated class or stable helper API.

Compilation cost and longer warmup
---------------------------------
linux/build-times/results.json and the twelve .log files record three complete
builds per arm. Each starts with clean application outputs, while compiler and
library dependencies stay cached. They run offline, in rotating arm order, on
all four CPUs of the same 8 GiB Linux VM, with no HTTP load running concurrently.
CN1 debug means the JVM development mode, not an unoptimized native executable.
The JavaSE backend runtime, shared classes and Bench handler are compiled with
JDK 25, as in the repository JavaSE workflow. Native compilation uses cached JDK 8
and JavaAPI/translator classes, generates C from that same Bench source, then
uses the packaged Clang -O3 glibc link path. Spring uses Maven clean package;
its native arm adds the native profile and native:compile. These paths have
different tools and runtime surfaces, not just different javac flags. Framework
source recompilation in CN1's repository workflow is included; prebuilt toolchain
and JavaAPI preparation is excluded. Package downloads and tests are excluded.
An initial offline attempt found the Maven clean plugin uncached. It was downloaded
before the final, complete three-round series; that failure is not a timing sample.

Reproducing build timing needs the pinned CN1 checkout with its translator and
JavaAPI dependencies built, plus the toolchain directories used in build_bench.py.
Run:

python3 linux/prepare_build_bench.py --repo /path/to/CodenameOne --output-dir /path/to/build-input

This assembles that checkout subset and its cached ASM jars into repo.tar. Copy repo.tar and linux/build_bench.py into the benchmark
container and run python3 /build_bench.py. Install
Temurin 8 under /opt/jdk8, JDK 25 under /opt/jdk25, GraalVM under /opt/java. Java 8
and Java 25 installation versions are recorded in build-times/toolchains.txt.
The script uses /repo for the source snapshot and /src for emitted C, with the
builder image's /usr/local/bin/link.sh and /out for artifacts. The Spring project
and cached Maven dependencies stay at /spring and /m2. No container startup time
is part of the measurements. These scripts are reproduction aids from this local
setup, not portable one-command installers for every environment.

linux/warmed-pair repeats CN1 native and Spring/JDK 25 with identical 60-second
warmups PER ROUTE, then three consecutive 10-second measurement windows. This is
one fresh process per runtime, plaintext then JSON. CPUs 0,1 are for the server;
2,3 for wrk, two load threads and 32 connections. Body validation precedes load;
any socket or non-2xx error rejects the run. These new paired-window data are
separate from the earlier long-warmup JVM diagnostic and the 27-process matrix.
The warmup script waits for all compilation measurements to finish first. A single
jcmd Compiler.codelist snapshot was taken during the JVM warmup (verified wrk -d60s
was running), outside the measured windows. It records tier-4 compiled methods
for DispatcherServlet.doDispatch and RequestMappingHandlerAdapter.invokeHandlerMethod.
The snapshot is retained as jit-codelist-during-warmup.txt.

runtime-baseline.json is the unchanged repository performance reference, not a
fresh benchmark. Each number is the median of calibration-run medians. 'runs'
counts calibration runs, not individual process pairs. plot_runtime_baselines.py
renders all sixteen self-translation configurations and all thirteen Linux ARM64
workloads, with the JDK as 1.0 and lower being better. Run it from any directory;
its charts go to charts/runtime in this bundle.

The Gradle file counts come from the tracked HelloCodenameOne sample at the pinned
commit, copied to a temporary directory and passed through the compiled current
GradleConversion implementation. gradle-conversion-counts.json lists every path:
358 files / 9 POMs become 305 files / 2 Kotlin Gradle configuration files. Build
output and caches are absent. Run:

python3 count_gradle_conversion.py --repo /path/to/CodenameOne \
  --output-dir /path/to/new-count-directory --java-home /path/to/jdk25

The checkout must have the build-engine compiled. The script reads the
BackendBeansTest Surefire classpath by default; use --classpath to supply it
explicitly. It retains the converter's original 8.0-SNAPSHOT version argument to
reproduce the recorded file count; the article's usable project example selects
the published 7.0.274 release instead.

