← All release test reports

AFM 0.9.19 · release evidence

Cross-model release qualification

PublishedSeptember 8, 2026Qualified targetQwen Next · DeepSeek V4 0731 · GLM 5.3 Flash · DwarfStar DSpark · Qwen 3.8 27B

Final v0.9.19 evidence covering six native MLX target and MTP modes, DwarfStar DSpark, Context through 32K, llmprobe conformance, and installed Qwen 3.8 27B testing on M3 Ultra and M4 Pro.

Apple M3 Ultra evidenceBrowse all 15 PNG charts and 7 HTML reports ↓

Qualification results

Native MLX MUST1,166 / 1,166

All llmprobe MUST assertions passed across six target and MTP modes.

Context coverage7 × through 32K

Six native MLX modes plus DwarfStar DSpark completed every context size.

Qwen comprehensive89 / 91 per host

Installed Qwen 3.8 27B runs on both M3 Ultra and M4 Pro; Codex judged all 91 records.

Live assertions116 per host

Both Qwen hosts also recorded two capability skips.

Promptfoo377 / 406 per host

76/76 native, 88/99 model behavior, and 213/231 forced-parser checks.

DwarfStar MUST179 / 184

97.3%; five known API limitations remain explicit.

Interpretation

Known caveats stay visible.

  • Each installed Qwen 3.8 27B host recorded 89/91 comprehensive assertion passes. The complete reports and Codex assessments preserve the two misses.
  • DwarfStar DSpark retains five API limitations: strict structured output in chat and Responses, chat logprobs, Messages stop-sequence reporting, and Messages token counting.
  • Speculation gains are workload-dependent. Qwen Next MTP was slower in the Context workload despite improving the separate llmprobe speculation scenario.
  • Host-memory charts describe total host memory rather than memory attributed only to the AFM process.

Bundle inventory

331 evidence files.

  • Installed Qwen 3.8 27B reports, assertions, Promptfoo results, and Codex assessments from M3 Ultra and M4 Pro
  • Six native MLX target/MTP reports with Context, llmprobe, server logs, and raw case data
  • DwarfStar DSpark Context and llmprobe evidence with its five known limitations
  • Consolidated and per-case throughput, TTFT, and host-memory charts with provenance metadata
  • Portable SHA-256 inventory for every curated evidence file

Hosted evidence · Apple M3 Ultra

M3 Ultra consolidated qualification

The six native MLX cases completed llmprobe and Context through 32K. DwarfStar DSpark also completed Context through 32K while retaining the five API limitations documented above.

Qualified build
v0.9.19-next.20260908.97a9acf
Run window
September 8, 2026 · 07:16–09:31 America/Toronto

Annotated consolidated charts

The final consolidated performance charts with visible k-label annotations.

Consolidated Context results

Cross-case Context charts for prompt speed, decode speed, TTFT, and whole-host memory.

Per-case Context and llmprobe evidence

A directly addressable Context chart and complete llmprobe HTML report for every qualified mode.

Canonical evidence bundle

afm-v0.9.19-test-reports.tar.gz

22.67 MiB compressed archive attached to the AFM 0.9.19 GitHub Release.

SHA-256
8e9194f99e518ad24e64a5fcbbec5f7922b5ea36e1fd692f2d9e86c7934c0e84
Reference URL
https://maclocal.ai/test-reports/0.9.19