On two identifier-anonymized JavaScript inputs, FlowName used fewer API calls and total tokens. An automated, original-source-aware judge preferred more FlowName names in both runs. These are individual case studies, not a general benchmark guarantee.
| Input | Bindings | API calls FlowName / Humanify | Total tokens FlowName / Humanify | Jev exclusive preferences FlowName / Humanify |
|---|---|---|---|---|
| Cigar | 411 | 65 / 411 | 87,107 / 137,332 | 155 / 37 |
| GrobPaint | 1,786 | 408 / 1,786 | 386,416 / 547,006 | 512 / 274 |
Each input was anonymized by replacing eligible lexical binding names. Both naming tools received the same anonymized input and used gemini-3.5-flash-lite. Cigar contains 1,756 code lines and 411 targets. GrobPaint contains 5,860 code lines and 1,786 targets, built from the application's five JavaScript modules; the external JSZip dependency was excluded. The GrobPaint source comes from this MIT-licensed revision. The GrobPaint FlowName run used a local build with post-1.0.0 context changes.
Jev 1.13 separately judged whether each applied name was useful and compared the pair when both were useful. It received the original author name and bounded original-code excerpts; the naming tools did not receive the author names. Candidate order varied by binding and tool identities were hidden from Jev. Its scores are uncalibrated model outputs, not probabilities that a name is correct.
Read the quality result with care. Humanify's naming prompt did not contain the target spelling for 200 Cigar bindings and 154 GrobPaint bindings. Among targets present in its prompt, exclusive Jev preferences were 38:37 on Cigar and 384:274 on GrobPaint. The Cigar comparison therefore does not isolate the effect of FlowName's grouping. Jev also evaluated final applied names: collision-safe suffixes such as Layer2 can lower a score even when the original model proposal was Layer.
These are one run per tool per input, with different request strategies. Automated labels need human review, especially near decision boundaries. No runtime behavior-equivalence test was performed on the renamed outputs. The published detail pages show bounded excerpts from the anonymized input; raw provider requests and original named sources are kept in local audit files.