[parquet] Avoid timestamp predicate pushdown for incompatible file schemas#8797
Merged
Merged
Conversation
Contributor
Author
|
Close #8801 |
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Purpose
Prevent incorrect query results when the timestamp logical type in a Parquet
file is incompatible with the current Paimon field type.
Validate the Parquet timestamp unit and adjustedToUTC flag before converting
Paimon predicates. Skip predicate pushdown when the file uses a different
timestamp unit or UTC adjustment. Also extract the common primitive field
lookup logic for reuse by decimal and timestamp validation.
Close #8801
Tests
Add ParquetFiltersTest#testTimestampFileUnitMismatchCannotPushDown to verify
that predicates are not pushed down when the Parquet file and read type use
different MILLIS/MICROS units.
Existing timestamp tests verify that predicates are still pushed down when
the Parquet timestamp schema matches the read type.
Manually verified with Spark SQL that timestamp equality queries return the
expected row after applying the fix.