Skip to content

Support decoding ints outside int64/uint64 - #469

Merged
jcrist merged 1 commit into
mainfrom
json-bigint
Jul 6, 2023
Merged

jcrist merged 1 commit into
mainfrom
json-bigint

Conversation

@jcrist

@jcrist jcrist commented Jul 6, 2023

Copy link
Copy Markdown
Member

Previously all decode methods would only support integers that fit in a uint64/int64. This PR removes that restriction for everything except MessagePack which as a binary protocol cannot support arbitrarily large integers.

To handle the parsing of these "big ints" we use CPython's builtin str -> int routine. This is slower than our custom str -> int function, but since integers of this scale are rare this should be fine in practice.

Note that currently CPython's str -> int routine exhibits quadratic behavior in the length of its input. This led to a DDOS CVE which was mitigated in 3.11 by adding a configurable max str length for int parsing. We hardcode the limit of 4300 chars here to support older python versions, while still properly handling cases where the user may have reduced this limit further. I think relying on this routine is fine, but if it becomes a problem later on we can rethink and maybe implement our own parsing routine. Parsing bigints will always be slower than parsing int64/uint64s, but there are better algorithms than the one currently used by CPython for handling this behavior.

BREAKING CHANGE: previously parsing a JSON "integer" (a number with no decimal or exponent components) into an Any or int | float field large would convert to a float if it didn't fit into an int64 or uint64. This is no longer the case. Unless a float is explicitly requested, a JSON number with no decimal or exponent components will always be parsed as an integer.

Previously all `decode` methods would only support integers that fit in
a `uint64`/`int64`. This PR removes that restriction for everything
_except_ MessagePack which as a binary protocol cannot support
arbitrarily large integers.

To handle the parsing of these "big ints" we use CPython's builtin `str
-> int` routine. This is slower than our custom `str -> int` function,
but since integers of this scale are rare this should be fine in
practice.

Note that currently CPython's `str -> int` routine exhibits quadratic
behavior in the length of its input. This led to a DDOS CVE which was
mitigated in 3.11 by adding a configurable max str length for int
parsing. We hardcode the limit of 4300 chars here to support older
python versions, while still properly handling cases where the user may
have reduced this limit further. I think relying on this routine is
fine, but if it becomes a problem later on we can rethink and maybe
implement our own parsing routine. Parsing bigints will always be slower
than parsing int64/uint64s, but there are better algorithms than the one
currently used by CPython for handling this behavior.

BREAKING CHANGE: previously parsing a JSON "integer" (a number with no
decimal or exponent components) into an `Any` or `int | float` field
large would convert to a float if it didn't fit into an `int64` or
`uint64`. This is no longer the case. Unless a `float` is explicitly
requested, a JSON number with no decimal or exponent components will
always be parsed as an integer.
@jcrist
jcrist merged commit 73582ae into main Jul 6, 2023
@jcrist
jcrist deleted the json-bigint branch July 6, 2023 01:25
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant