Translate format: regexps into patterns valid under Ajv's u flag - #66
Conversation
The Ruby -> ECMA-262 translation emitted identity escapes Unicode mode rejects (\-, \#, \ ), so everyday regexps and the :email/:url presets published patterns Ajv throws on (#47). It also exported \s, . and backreferences with looser ECMA-262 meanings, let {,n}, && and (?>...) through, and rewrote \A/\z with a gsub that corrupted \\A. Translation is now one whitelist tokenizer pass that reads escape pairs and character classes as units; anything it cannot carry exactly exports as x-permittable-pattern. Presets no longer need to skip a scan. A new spec compiles every exported pattern with node under the u flag and checks it agrees with Ruby on edge-case samples. Closes #47 Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
# Conflicts: # CHANGELOG.md
Review of #66: - Ruby's \b counts non-ASCII letters as word characters and Unicode mode's does not, so /\Aa\b/ exported a looser rule. \b and \B are refused outside a class; [\b] (backspace) still translates. - && or [ straight after a range hyphen slipped past the class check: [$-&&%] matches nothing in Ruby and "%" in ECMA-262. - [\s\S], exported on master, was refused. A class with both \s and \S now exports as [\s\S] (or [^\s\S]), and a lone [\S] as the complement of Ruby's \s. - The \u spec case now actually reaches the translation as \u{41}. - The node spec's cases live in a module instead of top-level constants. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
|
Review fixes pushed in bfc9812, a fast-forward on top of a merge of
655 examples, 0 failures, including the node spec. RuboCop is clean. A wider fuzz ran 80,000 translated regexps with 0 compile errors and 0 disagreements. The same harness finds 105 🤖 Generated with Claude Code |
optional :x, :float, in: (10**400..) dropped `minimum` entirely: the inward-nudge loop asks the field's own cast whether the published bound is honoured, and that cast runs the bound through `to_f`, which overflows a value this large to Infinity and walked the bound to Infinity too, where it was omitted like a genuinely infinite one. Master published it outright as `minimum: 10**400`, so this was a regression, not a labelled divergence: a JSON integer has no size limit, so an integral bound needs neither Float conversion nor the nudge machinery, which exists only to protect a FRACTIONAL bound from double-rounding. `json_bound` now returns an integral bound before the loop runs whenever `to_f` would overflow it. A genuinely fractional bound past Float::MAX still has no arbitrary- precision JSON representation to fall back on and stays omitted, now documented as deliberate rather than silent. Also, from the last regression pass: - Qualify the "published range can only be narrower than the enforced one" claim in the CHANGELOG/README: bounds are exact as published, but a client parsing with ordinary double-precision floats can still round a value across the boundary. That is inherent to double parsing, not an exporter bug, so no code change. - Make the CHANGELOG/README looser-divergence counts match what spec/schema_conformance_spec.rb actually asserts (LOOSER has six keys, not seven): split the "rules only an extension can carry" bullet so each of the six maps to one bullet, and say plainly that an untranslatable format: regexp has no conformance case yet, since PR #66 owns the regexp-translation rework it would depend on. - Delete the stale "No behaviour changes: this release adds tests and documentation only" line — Unreleased now contains substantial behaviour changes, this bound fix included. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
char_class refused any hyphen immediately following a completed range
the same way it refused one beside a class escape (\d/\w/\s), on the
theory that [a-c-e] means a different range under the u flag. It does
not: both Ruby and ECMA-262 read that hyphen as an ordinary member,
so [a-c-e] is the set {a, b, c, -, e} in both. A format: as common as
/\A[a-zA-Z0-9-_]+\z/ therefore exported no pattern at all, where
origin/master translated it correctly - a regression this PR
introduced while refusing && after a range hyphen.
The hyphen is now carried as a literal class member (and may itself
reopen a further range, chaining exactly as Ruby's own parser reads a
repeated hyphen, [a-z--x]). The class-escape refusal stays, verified
against real Ruby: [\d-z] and [a-\d] raise RegexpError building the
source at all, on either side, so that combination can never reach
the translator - the refusal is a defensive backstop, not something a
live disagreement between the two engines depends on.
Also corrects two doc/behavior mismatches found in the same pass:
- CHANGELOG said "\S beside anything else is refused"; the code
actually collapses a class holding both \s and \S (plus any other
members) to [\s\S]/[^\s\S], since the pair already covers every
character. Only \S without \s is refused when anything else shares
the class.
- README's fallback rationale implied every refusal is either a
Ruby-only construct or one ECMA-262 reads differently, which no
longer holds once the over-refusal above is fixed; reworded to name
the third case (a fragment with no ECMA-262 spelling once folded
into the rest of its class) and to call out the hyphen fix by name.
spec/json_schema_spec.rb gains a failing-first regression case for
the corpus in the finding ([a-zA-Z0-9-_], [A-Za-z0-9-_.], etc.) plus
a case documenting the class-escape source is unconstructable. The
node-verification spec's corpus gains the same cases, each checked to
compile under node -u and agree with Ruby on samples that include the
critical negative ("d" for [a-c-e]).
Re-ran the fuzz harness from the previous round (fuzz_F66.rb) after
the fix: 16 seeds, lengths 1-14, 80,000 translated regexps, 0 compile
errors under node -u, 0 disagreements with Ruby - same scale as the
prior round, still clean.
657 examples, 0 failures. RuboCop clean.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Round 2 (final regression pass)A regression review at bfc9812 found that fixing " Fix: that hyphen is now carried as a literal class member, and may itself reopen a further range ( Also fixed two doc mismatches found in the same pass:
Process: TDD — added a failing spec reproducing Verification:
Commit: caaf6f5 🤖 Generated with Claude Code |
…ce (#60) * Export decimal bounds as numbers and label two looser divergences A :decimal bounded by BigDecimals published "minimum": "0.01" as a string, because range bounds went through the authored-value re-encoding; the metaschema requires numbers, so the document was invalid. Bounds are now Integers when exact, otherwise Floats rounded inward when a double cannot round-trip the bound, so the published range is never wider than the enforced one. normalize: is now exported as x-permittable-normalize, and it and the string encoding of a bounded :decimal join max_depth: as documented, direction-asserted divergences in the README and the conformance spec. No runtime behaviour changes. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * Verify bounds against the field's cast and complete the looser list An infinite or NaN range endpoint is now omitted instead of crashing the export (BigDecimal("Infinity").to_i raised) or emitting Infinity. The inward nudge for an unrepresentable bound now asks the field's own cast and comparison, so a :float bound is safe too; it used to be one double short there. The documented looser divergences now also cover strings whose validity is only a format annotation (:decimal, :date, :datetime), validate: procs, and non-numeric ranges. The conformance spec checks each marker on the diverging field's own schema, and its tables live in a module. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * Publish an integral bound exactly, at any magnitude past Float::MAX optional :x, :float, in: (10**400..) dropped `minimum` entirely: the inward-nudge loop asks the field's own cast whether the published bound is honoured, and that cast runs the bound through `to_f`, which overflows a value this large to Infinity and walked the bound to Infinity too, where it was omitted like a genuinely infinite one. Master published it outright as `minimum: 10**400`, so this was a regression, not a labelled divergence: a JSON integer has no size limit, so an integral bound needs neither Float conversion nor the nudge machinery, which exists only to protect a FRACTIONAL bound from double-rounding. `json_bound` now returns an integral bound before the loop runs whenever `to_f` would overflow it. A genuinely fractional bound past Float::MAX still has no arbitrary- precision JSON representation to fall back on and stays omitted, now documented as deliberate rather than silent. Also, from the last regression pass: - Qualify the "published range can only be narrower than the enforced one" claim in the CHANGELOG/README: bounds are exact as published, but a client parsing with ordinary double-precision floats can still round a value across the boundary. That is inherent to double parsing, not an exporter bug, so no code change. - Make the CHANGELOG/README looser-divergence counts match what spec/schema_conformance_spec.rb actually asserts (LOOSER has six keys, not seven): split the "rules only an extension can carry" bullet so each of the six maps to one bullet, and say plainly that an untranslatable format: regexp has no conformance case yet, since PR #66 owns the regexp-translation rework it would depend on. - Delete the stale "No behaviour changes: this release adds tests and documentation only" line — Unreleased now contains substantial behaviour changes, this bound fix included. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> --------- Co-authored-by: Sang <sang@Soangs-MacBook-Pro.local> Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
CHANGELOG only: keep master's entries ahead of this PR's. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Master had moved to include #59-#61, #64 and #66 since this branch's last merge, plus #60 and #65 which both touch json_schema.rb's :decimal/:datetime export and lib/permittable/rspec.rb's `in:`/`default:` matcher chains — real conflicts, not just additive. lib/permittable.rb: kept this branch's permittable_transform seam (AuthoredValues needs it to walk a default without running transform: on it) — same guard (`violations.length == before`) as master's inline form. lib/permittable/json_schema.rb (4 hunks): combined rather than picked a side. - apply_in!'s enum export now uses BOTH #65's `field[:in_published] || allowed` (keeps a :date/:datetime member's authored String form) AND this branch's `decimal: :number` (keeps a :decimal member numeric and consistent with its own default/example export). - kept #60's `apply_normalize!` method (unrelated to this branch) ahead of json_value, whose signature this branch changes to `decimal: :string`. - json_value's Hash/BigDecimal/Time cases: kept this branch's `decimal:` threading for Hash/BigDecimal, and #65's `exact_iso8601` for Time (fixes sub-second Time :in members exporting as whole seconds — this branch's plain `.utc.iso8601` would have regressed that). - kept both decimal_json (this branch) and exact_iso8601 (#65) methods; each is used from the merged json_value body above. Verified with a probe: a :decimal in: list now exports numerically AND agrees with its own numeric default, and a :datetime in: list with sub-second members still exports the authored string via in_published — both at once, which neither branch alone tested. lib/permittable/rspec.rb (2 hunks) and spec/matchers_spec.rb: purely additive — #65's `cast_in`/`same_in?`/`within` alongside this branch's `default_mismatch`/`cast_default`. Combined by keeping both. README.md: combined the `default:` row (this branch, the transform: exception) with the `validate:` row (master, the array-skip-on-failed- validate note) — same table, different rows each PR had touched. Fixed along the way (found while merging, not part of either PR): - spec/permittable_spec.rb: a #61 perf spec asserted `valid_encoding?` is called on the caller's OWN string object, but this branch's permittable_own copies a request's String before Coercion.cast ever sees it (to avoid aliasing params) — so the original object is never touched, though the copy is (correctly) scanned exactly once. Rewrote the spec to count via a shared counter that survives `dup`, so it asserts the real invariant (one scan overall) rather than one specific object's identity. - spec/matchers_spec.rb: my own conflict resolution (both PRs inserted a new `it` block at the same point in the file) left one block's closing `end` missing — a syntax error that silently dropped all 42 examples in this file from the suite total without failing the run. Fixed; the suite total is now 819 (was silently 777). Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Bumps version.rb, retitles the Unreleased CHANGELOG section, and includes the regenerated Gemfile.lock — CI runs bundler in frozen mode, so a version bump without the lockfile fails the tag build. Twenty PRs since 0.8.0 (#16, #35, #51-#57, #59-#66, #71-#73): Rails 8 params.expect support and model-aware drafting in permittable:generate, the permittable:audit coverage command, reusable field groups (Permittable.fields/use), accept_params/reject_params RSpec matchers, RFC 9457 problem+json, named format: presets, and a type-checking mode for the schema-drift guard — plus a from-scratch Ruby-to-ECMA-262 pattern translator, a correctness overhaul of in: list casting and authored default:/example: storage, several encoding-crash and log-forging fixes, and three freshly-found gaps: an array field's required: + default: silently behaved as optional, accept_params/reject_params disagreed with a standalone Contract's own unknown: strictness, and the generator could draft a field from a permit call that only existed inside a log string. Minor rather than patch: mostly new surface, but a :string field's in: list of Symbols now matches correctly where it used to reject every request, and an array field combining required: and default: now fails at class load instead of silently treating the field as optional. See CHANGELOG.md for the complete list. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
The bug
Ajv compiles every
patternasnew RegExp(pattern, "u"), and Unicode mode is much stricter than Ruby. The translation copied Ruby's source and swapped\A/\zfor^/$, so:Every one of these is a real
patternin the published document, not anx-permittable-patternfallback.The fix: one tokenizer pass that only lets through what it knows
JsonSchema::EcmaPattern(new file) reads the source one token at a time. Escape pairs and character classes are read as whole units. Every token is either rewritten into something both engines read the same way, or the whole regexp exports asx-permittable-pattern:\A/\z^/$(a real anchor only, so\\Astays a backslash and an A)\s/\S[ \t\n\v\f\r]/[^ \t\n\v\f\r], and[^@\s]becomes[^@ \t\n\v\f\r][\S]/[\s\S][^ \t\n\v\f\r]/[\s\S](together they cover every character, whatever\smeans).\Sbeside anything else in a class is refused[\b].[^\n], because ECMA-262's.also stops at\r, U+2028 and U+2029\-outside a class,\#,\,\_\-inside a class, and a bare-right after a completed range ([a-c-e]){/}that Ruby reads as a literal\{/\}(?>…)(?#…)(?'n'…){,n}{n}?&&nested classes, octal,\e\x7\g<>, backreferences,\b/\Boutside a class,&&or[straight after a range hyphen, a stacked quantifier, a quantified lookaround, a repeated group namex-permittable-patternIt is a whitelist on purpose. If nobody thought of a construct, it is refused rather than published, and that keeps the old rule: a wrong pattern in published docs is worse than a missing one. The comment on the module lists each refusal and the reason for it.
The decisions I'd most like reviewed
1. Presets no longer skip the check. They used to skip the regexp scan (
vouched:), because that scan read the*+inside:email's character class as a possessive quantifier. The tokenizer reads a class as a unit, so the skip is no longer needed and I removed it. Presets now take the same path as an app's own regexp. That shared path is also what fixes #47::email's\#is written as a plain#when it is exported, and the definition does not change. It is stillURI::MailTo::EMAIL_REGEXPitself.2. Backreferences are now refused. You asked me to decide between
\1as a backreference and\1as octal. I went further and refused both. A backreference to a group that took no part in the match fails in Ruby, but in ECMA-262 it matches the empty string. So/\A(a)?\1b\z/rejects"b"on the server while its exported pattern accepts it, which is the unsafe direction. Also,\12can be a backreference or an octal escape depending on how many groups come before it. Backreferences in validation regexps are rare, and before this change they exported with that looser meaning.3.
.→[^\n]is beyond what the issue asked for. It is the same kind of bug as\s. Ruby's.accepts\rand U+2028 and ECMA-262's does not, so the old export was stricter than the server. The translation now matches Ruby exactly.4. Golden fixture. The email pattern in
spec/fixtures/openapi.jsonchanges from^[^@\\s]+@[^@\\s]+$to^[^@ \\t\\n\\v\\f\\r]+@[^@ \\t\\n\\v\\f\\r]+$. That is the\sfix and nothing else.^\\d{5}$does not change.Also
\p{...}is refused, because Ruby's property names differ from ECMA-262's. In practice these regexps never reach the translation: Ruby marks them encoding-fixed, and the existingoptions.zero?guard already exports them asx-permittable-pattern. That was true before this PR too.Review round 1
\b/\Bare refused outside a class. Ruby's word boundary countséas a word character (although its\wdoes not). Unicode mode's is ASCII-only, so/\Aa\b/rejects"aé"on the server while^a\baccepts it.\Bis off the other way. This was also exported as-is on master.[\b](backspace) still translates.&&or[straight after a range hyphen is refused.[$-&&%]matches nothing in Ruby but accepts%in ECMA-262, and[a-&&z]is a SyntaxError in Unicode mode.[\s\S]exports again. This was a regression against master. A class with both\sand\Sexports as[\s\S](or[^\s\S]), and a lone[\S]/[^\S]becomes the exact complement of Ruby's\s.\uspec case now really reaches the translation. Ruby rewrites\u0041toAin the source, even throughRegexp.new, so the case usesRegexp.new('\A\u{41}\x42\z').EcmaPatternSpecmodule instead of top-level constants.origin/masteris merged in. The golden fixture's email pattern for the new master operations now uses the spelled-out\stoo.]case and the new hyphen cases are built with warnings off, so the suite prints no regexp warnings.Review round 2 (final regression pass)
-right after a completed range was refused — a regression this PR itself introduced. Fixing "&&or[straight after a range hyphen" (round 1) over-tightened the same guard: it treated "the previous member was a range" the same as "the previous member was a class escape," and refused both. It shouldn't have.[a-c-e]is not "a different range in Unicode mode" — both Ruby and ECMA-262 read that hyphen as an ordinary member once a range has just closed, so[a-c-e]is the set{a, b, c, -, e}(rejecting"d") in both. Aformat:as ordinary as/\A[a-zA-Z0-9-_]+\z/therefore published no pattern at all, whereorigin/mastertranslated it correctly. The hyphen is now carried as a literal member, and may itself reopen a further range ([a-z--x], which is exactly how Ruby's own parser reads a second hyphen). The class-escape half of the guard stays, but I checked it against real Ruby first rather than assuming:Regexp.new('[\d-z]')andRegexp.new('[a-\d]')both raiseRegexpError— a class escape can never sit beside a range hyphen in a real Regexp, on either side — so that refusal is a defensive backstop, not something a live disagreement between the two engines depends on.\Sbeside anything else is refused," but the code actually collapses a class holding both\sand\S— plus any other members — to[\s\S]/[^\s\S], since the pair already covers every character whatever either escape means; only\Swithout\sis refused when anything else shares the class. And the README's fallback rationale implied every refusal is either a Ruby-only construct or one ECMA-262 reads differently, which stopped being true once the over-refusal above needed fixing — there's a third case, a fragment with no ECMA-262 spelling once folded into the rest of its own class (\Sbeside a member other than\s). Both are reworded to state the actual rules.spec/json_schema_spec.rb), watched it fail against the pre-fix code, then fixedchar_class. The node-verification spec's corpus (spec/ecma_pattern_spec.rb) gains the same cases plus the chained-hyphen case, each checked to compile undernode -uand to agree with Ruby on samples that include the negative case ("d"for[a-c-e]).Verification
spec/ecma_pattern_spec.rbruns every exported pattern through node with theuflag in a single process: every preset, a table of app-style regexps, and everypatternin the golden fixture. It also checks that each translated pattern accepts exactly what Ruby accepts on edge-case samples: NBSP, U+FEFF, U+2028 and U+3000 for\s;\rand U+2028 for.;#inside:email's local part;\\A;{5}; and now the hyphen-after-range corpus. Ifnodeis not on PATH it skips with a message. GitHub's runners have node, so it runs in CI.uand 0 disagreed with Ruby on the samples. Before that decision it found exactly one disagreement, the unset-backreference case above.\b,\B,\W,[\s\S],$-&,\u{e9}and\&tokens andé,éa,%, NUL and an emoji as samples. Across 80,000 translated regexps (16 seeds, lengths 1–14): 0 compile errors and 0 disagreements with Ruby. Pointed at the pre-review tokenizer, the same harness finds 105\b/\Bdisagreements on one seed, so it does catch that class of bug.fuzz_F66.rb, unchanged) against the fixed code: same 16 seeds, lengths 1–14, 80,000 translated regexps, 0 compile errors, 0 disagreements with Ruby — still clean at the same scale.Closes #47
🤖 Generated with Claude Code