Support decoding ints outside int64/uint64 - #469
Merged
Merged
Conversation
Previously all `decode` methods would only support integers that fit in a `uint64`/`int64`. This PR removes that restriction for everything _except_ MessagePack which as a binary protocol cannot support arbitrarily large integers. To handle the parsing of these "big ints" we use CPython's builtin `str -> int` routine. This is slower than our custom `str -> int` function, but since integers of this scale are rare this should be fine in practice. Note that currently CPython's `str -> int` routine exhibits quadratic behavior in the length of its input. This led to a DDOS CVE which was mitigated in 3.11 by adding a configurable max str length for int parsing. We hardcode the limit of 4300 chars here to support older python versions, while still properly handling cases where the user may have reduced this limit further. I think relying on this routine is fine, but if it becomes a problem later on we can rethink and maybe implement our own parsing routine. Parsing bigints will always be slower than parsing int64/uint64s, but there are better algorithms than the one currently used by CPython for handling this behavior. BREAKING CHANGE: previously parsing a JSON "integer" (a number with no decimal or exponent components) into an `Any` or `int | float` field large would convert to a float if it didn't fit into an `int64` or `uint64`. This is no longer the case. Unless a `float` is explicitly requested, a JSON number with no decimal or exponent components will always be parsed as an integer.
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Previously all
decodemethods would only support integers that fit in auint64/int64. This PR removes that restriction for everything except MessagePack which as a binary protocol cannot support arbitrarily large integers.To handle the parsing of these "big ints" we use CPython's builtin
str -> introutine. This is slower than our customstr -> intfunction, but since integers of this scale are rare this should be fine in practice.Note that currently CPython's
str -> introutine exhibits quadratic behavior in the length of its input. This led to a DDOS CVE which was mitigated in 3.11 by adding a configurable max str length for int parsing. We hardcode the limit of 4300 chars here to support older python versions, while still properly handling cases where the user may have reduced this limit further. I think relying on this routine is fine, but if it becomes a problem later on we can rethink and maybe implement our own parsing routine. Parsing bigints will always be slower than parsing int64/uint64s, but there are better algorithms than the one currently used by CPython for handling this behavior.BREAKING CHANGE: previously parsing a JSON "integer" (a number with no decimal or exponent components) into an
Anyorint | floatfield large would convert to a float if it didn't fit into anint64oruint64. This is no longer the case. Unless afloatis explicitly requested, a JSON number with no decimal or exponent components will always be parsed as an integer.