package resolver
- Alphabetic
- By Inheritance
- resolver
- AnyRef
- Any
- Hide All
- Show All
- Public
- Protected
Type Members
- class AggregateExpressionResolver extends TreeNodeResolver[AggregateExpression, Expression] with ResolvesExpressionChildren with CoercesExpressionTypes
Resolver for AggregateExpressions that can come from either FunctionResolver or ExpressionResolver.
Resolver for AggregateExpressions that can come from either FunctionResolver or ExpressionResolver. It handles the resolution and validation of AggregateExpression.
- case class AggregateResolutionResult(operator: LogicalPlan, outputList: Seq[NamedExpression], groupingAttributeIds: HashSet[ExprId], aggregateListAliases: Seq[Alias], baseAggregate: Aggregate) extends Product with Serializable
Stores the resulting operator, output list, grouping attributes, list of aliases from aggregate list and base Aggregate, obtained by resolving an Aggregate operator.
- class AggregateResolver extends TreeNodeResolver[Aggregate, LogicalPlan] with AliasHelper
Resolves an Aggregate by resolving its child, aggregate expressions and grouping expressions.
Resolves an Aggregate by resolving its child, aggregate expressions and grouping expressions. Updates the NameScopeStack with its output and performs validation related to Aggregate resolution.
- case class AggregateWithLcaResolutionResult(resolvedOperator: LogicalPlan, outputList: Seq[NamedExpression], aggregateListAliases: Seq[Alias], baseAggregate: Aggregate) extends Product with Serializable
Stores the result of resolution of lateral column aliases in an Aggregate.
Stores the result of resolution of lateral column aliases in an Aggregate.
- resolvedOperator
The resolved operator.
- outputList
The output list of the resolved operator.
- aggregateListAliases
List of aliases from aggregate list and all artificially inserted Project nodes.
- baseAggregate
Aggregate node constructed by LateralColumnAliasResolver while resolving lateral column references in Aggregate.
- class AliasResolver extends TreeNodeResolver[UnresolvedAlias, Expression] with ResolvesExpressionChildren with AliasHelper
Resolver class that resolves unresolved aliases and handles user-specified aliases.
- case class AnalyzerBridgeState(relationsWithResolvedMetadata: RelationsWithResolvedMetadata = new AnalyzerBridgeState.RelationsWithResolvedMetadata, catalogRelationsWithResolvedMetadata: CatalogRelationsWithResolvedMetadata = new AnalyzerBridgeState.CatalogRelationsWithResolvedMetadata, hiveRelationsWithResolvedMetadata: HiveRelationsWithResolvedMetadata = new AnalyzerBridgeState.HiveRelationsWithResolvedMetadata) extends Product with Serializable
The AnalyzerBridgeState is a state passed from legacy Analyzer to the single-pass Resolver.
The AnalyzerBridgeState is a state passed from legacy Analyzer to the single-pass Resolver. It is used in dual-run mode (when ANALYZER_SINGLE_PASS_RESOLVER_RELATION_BRIDGING_ENABLED is true).
- relationsWithResolvedMetadata
A map from BridgedRelationId to the relations with resolved metadata. It allows us to reuse the relation metadata and avoid duplicate catalog/table lookups.
- catalogRelationsWithResolvedMetadata
A map from UnresolvedCatalogRelation to the relations with resolved metadata. It allows us to reuse the relation metadata and avoid duplicate catalog/table lookups.
- hiveRelationsWithResolvedMetadata
A map from HiveTableRelation to their resolved LogicalRelation counterparts. We cannot import those nodes here because of recursive dependencies, so we rely on overridden LogicalPlan.equals and LogicalPlan.hashCode. Keys are canonicalized to compensate for stats added by DetermineTableStats.
- case class AttributeScope(attributes: AttributeSet, isSubqueryRoot: Boolean = false) extends Product with Serializable
A scope with registered attributes encountered during the logical plan validation process.
A scope with registered attributes encountered during the logical plan validation process. We use AttributeSet here to check the equality of attributes based on their expression IDs.
- class AttributeScopeStack extends AnyRef
The AttributeScopeStack is used to validate that the attribute which was encountered by the ExpressionResolutionValidator is in the current operator's visibility scope.
The AttributeScopeStack is used to validate that the attribute which was encountered by the ExpressionResolutionValidator is in the current operator's visibility scope.
E.g. for the following SQL query:
SELECT a, a, a + col2 FROM (SELECT col1 as a, col2 FROM VALUES (1, 2));
Having the following logical plan:
Project [a#2, a#2, (a#2 + col2#1) AS (a + col2)#3] +- SubqueryAlias __auto_generated_subquery_name +- Project [col1#0 AS a#2, col2#1] +- LocalRelation [col1#0, col2#1]
The LocalRelation outputs attributes with IDs #0 and #1, which can be referenced by the lower Project. This Project produces a new attribute ID #2 for an alias and retains the old ID #1 for col2. The upper Project references
atwice using the same ID #2 and produces a new ID #3 for an alias ofa + col2. - class AutoGeneratedAliasProvider extends AnyRef
AutoGeneratedAliasProvider is a tool to create auto-generated aliases in the plan.
AutoGeneratedAliasProvider is a tool to create auto-generated aliases in the plan. All the auto-generated aliases have to be registered in ExpressionIdAssigner.
- class BinaryArithmeticResolver extends TreeNodeResolver[BinaryArithmetic, Expression] with ProducesUnresolvedSubtree with CoercesExpressionTypes
BinaryArithmeticResolver is invoked by ExpressionResolver in order to resolve BinaryArithmetic nodes.
BinaryArithmeticResolver is invoked by ExpressionResolver in order to resolve BinaryArithmetic nodes. During resolution, calling BinaryArithmeticWithDatetimeResolver and applying type coercion can result in BinaryArithmetic producing some other type of node or a subtree of nodes. In such cases a downwards traversal is necessary, but not going deeper than the original expression's children, since all nodes below that point are guaranteed to be already resolved.
For example, given a query:
SELECT '4 11:11' - INTERVAL '4 22:12' DAY TO MINUTE
BinaryArithmeticResolver is called for the following expression:
Subtract( Literal('4 11:11', StringType), Literal(Interval('4 22:12' DAY TO MINUTE), DayTimeIntervalType(0,2)) )
After calling BinaryArithmeticWithDatetimeResolver and applying type coercion, the expression is transformed into:
Cast( DatetimeSub( TimestampAddInterval( Literal('4 11:11', StringType), UnaryMinus( Literal(Interval('4 22:12' DAY TO MINUTE), DayTimeIntervalType(0,2)) ) ) ) )
A single Subtract node is replaced with a subtree of nodes. In order to resolve this subtree we need to invoke ExpressionResolver recursively on the top-most node's children. The top-most node itself is not resolved recursively in order to avoid recursive calls to BinaryArithmeticResolver and other sub-resolvers. To prevent a case where we resolve the same node twice, we need to mark nodes that will act as a limit for the downwards traversal by applying a ResolverTag.SINGLE_PASS_SUBTREE_BOUNDARY tag to them. These children along with all the nodes below them are guaranteed to be resolved at this point. When ExpressionResolver reaches one of the tagged nodes, it returns identity rather than resolving it. Finally, after resolving the subtree, we need to resolve the top-most node itself, which in this case means applying a timezone, if necessary.
- case class BridgedRelationId(unresolvedRelation: UnresolvedRelation, catalogAndNamespace: Seq[String]) extends Product with Serializable
The BridgedRelationId is a unique identifier for an unresolved relation in the whole logical plan including all the nested views.
The BridgedRelationId is a unique identifier for an unresolved relation in the whole logical plan including all the nested views. It is used to lookup relations with resolved metadata which were processed by the fixed-point when running two Analyzers in dual-run mode. Storing catalogAndNamespace is required to differentiate tables/views created in different catalogs as their UnresolvedRelations could have same structure.
- class BridgedRelationMetadataProvider extends RelationMetadataProvider
The BridgedRelationMetadataProvider is a RelationMetadataProvider that just reuses resolved metadata from the AnalyzerBridgeState.
The BridgedRelationMetadataProvider is a RelationMetadataProvider that just reuses resolved metadata from the AnalyzerBridgeState. This is used in the single-pass Resolver to avoid duplicate catalog/table lookups in dual-run mode, so metadata is simply reused from the fixed-point Analyzer run. We strictly rely on the AnalyzerBridgeState to avoid any blocking calls here.
- case class CandidatesForResolution(attributes: Seq[Attribute], outputType: OutputType) extends Product with Serializable
CandidatesForResolution is used by the NameScope during multipart name resolution to prioritize attributes from different types of operator output (main, hidden, metadata).
- trait CoercesExpressionTypes extends SQLConfHelper
CoercesExpressionTypes is extended by resolvers that need to apply type coercion.
CoercesExpressionTypes is extended by resolvers that need to apply type coercion.
ansiTransformationsandnonAnsiTransformationsmay be overridden with custom transformation lists. - class CteRegistry extends AnyRef
The CteRegistry is responsible for managing the stack of CteScopes and resolving visible CTERelationDef names.
- class CteScope extends AnyRef
The CteScope is responsible for keeping track of visible and known CTE definitions at a given stage of a SQL query/DataFrame program resolution.
The CteScope is responsible for keeping track of visible and known CTE definitions at a given stage of a SQL query/DataFrame program resolution. These scopes are stacked and the stack is managed by the CteRegistry. The scope is created per single WITH clause.
The CTE operators are:
- UnresolvedWith. This is a
hostoperator that contains a list of unresolved CTE definitions from the WITH clause and a single child operator, which is the actual unresolved SELECT query. - UnresolvedRelation. This is a generic unresolved relation operator that will sometimes be resolved to a CTE definition and later replaced with a CTERelationRef. The CTE takes precedence over a regular table or a view when resolving this identifier.
- CTERelationDef. This is a reusable logical plan, which will later be referenced by the lower CTE definitions and UnresolvedWith child.
- CTERelationRef. This is a leaf node similar to a relation operator that references a certain CTERelationDef by its ID. It has a name (unique locally for a WITH clause list) and an ID (unique for all the CTEs in a query).
- WithCTE. This is a
hostoperator that contains a list of resolved CTE definitions from the WITH clause and a single child operator, which is the actual resolved SELECT query.
The task of the Resolver is to correctly place WithCTE with CTERelationDefs inside and make sure that CTERelationRefs correctly reference CTERelationDefs with their IDs. The decision whether to inline those CTE subtrees or not is made by the Optimizer, unlike what Spark does for the Views (always inline during the analysis).
There are some caveats in how Spark places those operators and resolves their names:
- Ambiguous CTE definition names are disallowed only within a single WITH clause, and this is validated by the Parser in AstBuilder using QueryParsingErrors.duplicateCteDefinitionNamesError:
-- This is disallowed. WITH cte AS (SELECT 1), cte AS (SELECT 2) SELECT * FROM cte;
- When UnresolvedRelation identifier is resolved to a CTERelationDef and there is a name conflict on several layers of CTE definitions, the lower definitions take precedence:
-- The result is `3`, lower [[CTERelationDef]] takes precedence. WITH cte AS ( SELECT 1 ) SELECT * FROM ( WITH cte AS ( SELECT 2 ) SELECT * FROM ( WITH cte AS ( SELECT 3 ) SELECT * FROM cte ) )
- Any subquery can contain UnresolvedWith on top of it, but WithCTE is not gonna be 1 to 1 to its unresolved counterpart. For example, if we are dealing with simple subqueries, CTERelationDefs will be merged together under a single WithCTE. The previous example would produce the following resolved plan:
WithCTE :- CTERelationDef 18, false : +- ... :- CTERelationDef 19, false : +- ... :- CTERelationDef 20, false : +- ... +- Project [3#1203] : +- ...
- The WithCTE operator is placed on top of the resolved operator if one of the following
conditions are met:
- We just resolved an UnresolvedWith, which is the topmost UnresolvedWith of this root query, view or an expression subquery. 2. In case there is no single topmost UnresolvedWith, we pick the least common ancestor of those branches. This is going to be a multi-child operator - Union, Join, etc.
Here's an example for the second case:
SELECT * FROM ( WITH cte AS ( SELECT 1 ) SELECT * FROM cte UNION ALL ( WITH cte AS ( SELECT 2 ) SELECT * FROM cte ) )->
Project [1#60] +- SubqueryAlias __auto_generated_subquery_name +- WithCTE :- CTERelationDef 30, false : +- ... :- CTERelationDef 31, false : +- ... +- Union false, false :- Project [1#60] : +- ... +- Project [2#61] +- ...
Consider a different example though:
SELECT * FROM ( SELECT 1 UNION ALL ( WITH cte AS ( SELECT 2 ) SELECT * FROM cte ) )
The Union operator is not the least common ancestor of the UnresolvedWiths in the query. In fact, there's just a single UnresolvedWith, which is a proper place where we need to place a WithCTE.
- However, if we have any expression subquery (scalar/IN/EXISTS...), the top CTERelationDefs and subquery's CTERelationDef won't be merged together (as they are separated by an expression tree):
WITH cte AS ( SELECT 1 AS col1 ) SELECT * FROM cte WHERE col1 IN ( WITH cte AS ( SELECT 2 ) SELECT * FROM cte )
->
WithCTE :- CTERelationDef 21, false : +- ... +- Project [col1#1223] +- Filter col1#1223 IN (list#1222 []) : +- WithCTE : :- CTERelationDef 22, false : : +- ... : +- Project [2#1241] : +- ... +- ...
- Upper CTEs are visible through subqueries and can be referenced by lower operators, but not through the View boundary:
CREATE VIEW v1 AS SELECT 1; CREATE VIEW v2 AS SELECT * FROM v1; -- The result is 1. -- The `v2` body will be inlined in the main query tree during the analysis, but upper `v1` -- CTE definition _won't_ take precedence over the lower `v1` view. WITH v1 AS ( SELECT 2 ) SELECT * FROM v2;
- UnresolvedWith. This is a
- trait DelegatesResolutionToExtensions extends AnyRef
The DelegatesResolutionToExtensions is a trait which provides a method to delegate the resolution of unresolved operators to a list of ResolverExtensions.
- class ExplicitlyUnsupportedResolverFeature extends Exception
This is an addon to ResolverGuard functionality for features that cannot be determined by only looking at the unresolved plan.
This is an addon to ResolverGuard functionality for features that cannot be determined by only looking at the unresolved plan. Resolver will throw this control-flow exception when it encounters some explicitly unsupported feature. Later behavior depends on the value of HybridAnalyzer.exposeExplicitlyUnsupportedResolverFeature flag:
- If it is true: It will later be caught by HybridAnalyzer to abort single-pass
analysis without comparing single-pass and fixed-point results. The motivation for this
feature is the same as for the ResolverGuard - we want to have an explicit allowlist of
unimplemented features that we are aware of, and
UNSUPPORTED_SINGLE_PASS_ANALYZER_FEATUREwill signal us the rest of the gaps. - If it is false: It will be thrown by the HybridAnalyzer in order to get better sense of coverage.
For example, UnresolvedRelation can be intermediately resolved by ResolveRelations as UnresolvedCatalogRelation or a View (among all others). Say that for now the views are not implemented, and we are aware of that, so ExplicitlyUnsupportedResolverFeature will be thrown in the middle of the single-pass analysis to abort it.
- If it is true: It will later be caught by HybridAnalyzer to abort single-pass
analysis without comparing single-pass and fixed-point results. The motivation for this
feature is the same as for the ResolverGuard - we want to have an explicit allowlist of
unimplemented features that we are aware of, and
- class ExpressionIdAssigner extends AnyRef
ExpressionIdAssigner is used by the ExpressionResolver to assign unique expression IDs to NamedExpressions (AttributeReferences and Aliases).
ExpressionIdAssigner is used by the ExpressionResolver to assign unique expression IDs to NamedExpressions (AttributeReferences and Aliases). This is necessary to ensure that Optimizer performs its work correctly and does not produce correctness issues.
The framework works the following way:
- Each leaf operator must have globally unique output IDs (even if it's the same table, view, or CTE).
- The AttributeReferences get propagated "upwards" through the operator tree with their IDs preserved. In case of correlated subqueries AttributeReferences may propagate downwards from the outer scope to the point of correlated reference in the subquery. Currently only one level of correlation is supported.
- Each Alias gets assigned a new globally unique ID and it sticks with it after it gets converted to an AttributeReference when it is outputted from the operator that produced it.
- Any operator may have AttributeReferences with the same IDs in its output given it is the same attribute. Thus, **no multi-child operator may have children with conflicting AttributeReference IDs**. In other words, two subtrees must not output the AttributeReferences with the same IDs, since relations, views and CTEs all output unique attributes, and Aliases get assigned new IDs as well. ExpressionIdAssigner.assertOutputsHaveNoConflictingExpressionIds is used to assert this invariant.
For SQL queries, this framework provides correctness just by reallocating relation outputs and by validating the invariants mentioned above. Reallocation is done in Resolver.handleLeafOperator. If all the relations (even if it's the same table) have unique output IDs, the expression ID assignment will be correct, because there are no duplicate IDs in a pure unresolved tree. The old ID -> new ID mapping is not needed in this case. For example, consider this query:
SELECT * FROM t AS t1 CROSS JOIN t AS t2 ON t1.col1 = t2.col1
The analyzed plan should be:
Project [col1#0, col2#1, col1#2, col2#3] +- Join Cross, (col1#0 = col1#2) :- SubqueryAlias t1 : +- Relation t[col1#0,col2#1] parquet +- SubqueryAlias t2 +- Relation t[col1#2,col2#3] parquet
and not:
Project [col1#0, col2#1, col1#0, col2#1] +- Join Cross, (col1#0 = col1#0) :- SubqueryAlias t1 : +- Relation t[col1#0,col2#1] parquet +- SubqueryAlias t2 +- Relation t[col1#0,col2#1] parquet
Because in the latter case the join condition is always true.
For DataFrame programs we need the full power of ExpressionIdAssigner, and old ID -> new ID mapping comes in handy, because DataFrame programs pass _partially_ resolved plans to the Resolver, which may consist of duplicate subtrees, and thus will have already assigned expression IDs. These already resolved duplicate subtrees with assigned IDs will conflict. Hence, we need to reallocate all the leaf node outputs _and_ remap old IDs to the new ones. Also, DataFrame programs may introduce the same Aliases in different parts of the query plan, so we just reallocate all the Aliases.
For example, consider this DataFrame program:
spark.range(0, 10).select($"id").write.format("parquet").saveAsTable("t") val alias = ($"id" + 1).as("id") spark.table("t").select(alias).select(alias)
The analyzed plan should be:
Project [(id#6L + cast(1 as bigint)) AS id#13L] +- Project [(id#4L + cast(1 as bigint)) AS id#6L] +- SubqueryAlias spark_catalog.default.t +- Relation spark_catalog.default.t[id#4L] parquet
and not:
Project [(id#6L + cast(1 as bigint)) AS id#6L] +- Project [(id#4L + cast(1 as bigint)) AS id#6L] +- SubqueryAlias spark_catalog.default.t +- Relation spark_catalog.default.t[id#4L] parquet
Because the latter case will confuse the Optimizer and the top Project will be eliminated leading to incorrect result.
In case of partially resolved DataFrame subtrees with correlated subqueries inside we need to remap OuterReferences as well:
val df = spark.sql(""" SELECT * FROM t1 WHERE EXISTS ( SELECT * FROM t2 WHERE t2.id == t1.id ) """) df.union(df)
The analyzed plan should be:
Union false, false :- Project [id#1] : +- Filter exists#9 [id#1] : : +- Project [id#16] : : +- Filter (id#16 = outer(id#1)) : : +- SubqueryAlias spark_catalog.default.t2 : : +- Relation spark_catalog.default.t2[id#16] parquet : +- SubqueryAlias spark_catalog.default.t1 : +- Relation spark_catalog.default.t1[id#1] parquet +- Project [id#17 AS id#19] +- Project [id#17] +- Filter exists#9 [id#17] : +- Project [id#18] : +- Filter (id#18 = outer(id#17)) : +- SubqueryAlias spark_catalog.default.t2 : +- Relation spark_catalog.default.t2[id#18] parquet +- SubqueryAlias spark_catalog.default.t1 +- Relation spark_catalog.default.t1[id#17] parquet
Note how id#17 is the same in outer branch and in a subquery - is was properly remapped, because the right subtree of Union contained identical expression IDs as the left subtree. That's why we pass main mapping as outer mapping to the correlated subquery branch.
There's an important caveat here: those branches of a logical plan tree where outputs do not conflict. We should preserve expression IDs on those branches wherever possible because DataFrames may reference each other using their attributes. This also makes sense for performance reasons.
Consider this example:
val df1 = spark.range(0, 10).select($"id") val df2 = spark.range(5, 15).select($"id") df1.union(df2).filter(df1("id") === 5)
In this example
df("id")references loweridattribute by expression ID, sounionmust not reassign expression IDs indf1(left child). Referencingdf2(right child) is not supported in Spark, because Union does not output it, but we don't have to regenerate expression IDs in that branch either.However:
val df1 = spark.range(0, 10).select($"id") df1.union(df1).filter(df1("id") === 5)
Here we need to regenerate expression IDs in the right branch, because those would conflict (both branches are the same plan). Expression IDs in the left branch may be preserved.
CTE references are handled in a special way to stay compatible with the fixed-point Analyzer. First CTERelationRef that we meet in the query plan can preserve its output expression IDs, and the plan will be inlined by the InlineCTE without any artificial Aliases that "stitch" expression IDs together. This way we ensure that Optimizer behavior is the same as after the fixed-point Analyzer and that no extra projections are introduced.
The ExpressionIdAssigner covers both SQL and DataFrame scenarios with single approach and is integrated into the single-pass analysis framework.
The ExpressionIdAssigner is used in the following way:
- When the Resolver traverses the tree downwards prior to starting bottom-up analysis,
we build the mappingStack by calling pushMapping.
for every child of a multi-child operator, so we have a separate stack entry (separate
mapping) for each branch. This way sibling branches' mappings are isolated from each other and
attribute IDs are reused only within the same branch. Initially we push
None, because the mapping needs to be initialized later with the correct output of a resolved operator. - When the bottom-up analysis starts, we assign IDs to all the NamedExpressions which are
present in operators starting from the LeafNodes using mapExpression.
createMappingForLeafOperator is called right after each LeafNode is resolved, and
first remapped attributes come from that LeafNode. This is done if leaf operator output
doesn't conflict with
globalExpressionIds. - Once the child branch is resolved, a code block started with pushMapping ends by calling popMapping.
- After the multi-child operator is resolved, we call createMappingFromChildMappings to
initialize the mapping with attributes collected in popMapping with
collectChildMapping = true. - While traversing the expression tree, we may meet a SubqueryExpression and resolve its
plan. In this case we call pushMapping with
isSubqueryRoot = trueto pass the current mapping as outer mapping to the subquery branches. Any subquery branch may reference outer attributes, so ifisSubqueryRootisfalse, we pass the previousouterMappingto lower branches. Since we only support one level of correlation, for every subquery level currentmappingbecomesouterMappingfor the next level. - Continue remapping expressions until we reach the root of the operator tree.
- class ExpressionResolutionContext extends AnyRef
The ExpressionResolutionContext is a state that is propagated between the nodes of the expression tree during the bottom-up expression resolution process.
The ExpressionResolutionContext is a state that is propagated between the nodes of the expression tree during the bottom-up expression resolution process. This way we pass the results of ExpressionResolver.resolve call, which are not the resolved child itself, from children to parents.
- class ExpressionResolutionValidator extends AnyRef
The ExpressionResolutionValidator performs the validation work on the expression tree for the ResolutionValidator.
The ExpressionResolutionValidator performs the validation work on the expression tree for the ResolutionValidator. These two components work together recursively validating the logical plan. You can find more info in the ResolutionValidator scaladoc.
- class ExpressionResolver extends TreeNodeResolver[Expression, Expression] with ProducesUnresolvedSubtree with ResolvesExpressionChildren with CoercesExpressionTypes
The ExpressionResolver is used by the Resolver during the analysis to resolve expressions.
The ExpressionResolver is used by the Resolver during the analysis to resolve expressions.
The functions here generally traverse unresolved Expression nodes recursively, constructing and returning the resolved Expression nodes bottom-up. This is the primary entry point for implementing expression analysis, wherein the resolve method accepts a fully unresolved Expression and returns a fully resolved Expression in response with all data types and attribute reference ID assigned for valid requests. This resolver also takes responsibility to detect any errors in the initial SQL query or DataFrame and return appropriate error messages including precise parse locations wherever possible.
- case class ExpressionTreeTraversal(parentOperator: LogicalPlan, ansiMode: Boolean, lcaEnabled: Boolean, groupByAliases: Boolean, sessionLocalTimeZone: String, defaultCollation: Option[String] = None, invalidExpressionsInTheContextOfOperator: ArrayList[Expression] = new ArrayList[Expression], referencedAttributes: HashMap[ExprId, Attribute] = new HashMap[ExprId, Attribute], isFilterOnTopOfAggregate: Boolean = false, isSortOnTopOfAggregate: Boolean = false) extends Product with Serializable
Properties of a current expression tree traversal.
Properties of a current expression tree traversal. Settings like
ansiMode,lcaEnabledorgroupByAliasesare set once per expression tree traversal as an optimization to avoid frequent SQLConf lookups. These settings may be different for different views.- parentOperator
The parent operator of the current expression tree.
- ansiMode
Whether the current expression tree being resolved is in ANSI mode.
- lcaEnabled
Whether lateral column alias resolution is enabled.
- groupByAliases
Whether the group by aliases resolution is enabled.
- sessionLocalTimeZone
The session local time zone.
- defaultCollation
View's default collation if explicitly set.
- invalidExpressionsInTheContextOfOperator
The expressions that are invalid in the context of the current expression tree and its parent operator.
- referencedAttributes
All attributes that are referenced during the resolution of expression trees.
- isFilterOnTopOfAggregate
Whether the current expression tree is below a Filter on top of an Aggregate operator.
- isSortOnTopOfAggregate
Whether the current expression tree is below a Sort on top of an Aggregate operator (or on top of a Filter with an Aggregate as its child).
- class ExpressionTreeTraversalStack extends SQLConfHelper
The stack of expression tree traversal properties which are accumulated during the resolution of a certain expression tree.
The stack of expression tree traversal properties which are accumulated during the resolution of a certain expression tree. This is filled by the ExpressionResolver.resolveExpressionTreeInOperatorImpl, and will usually have size 1. However, in case of subquery expressions we would call ExpressionResolver.resolveExpressionTreeInOperatorImpl several times recursively for each expression tree in the operator tree -> expression tree -> operator tree -> expression tree -> ... chain. Consider this example:
SELECT col1 FROM VALUES (1) AS t1 WHERE EXISTS ( SELECT * FROM VALUES (2) AS t2 WHERE (SELECT col1 FROM VALUES (3) AS t3) == t1.col1 )
We would have 3 nested stack entries for while resolving the lower scalar subquery (with the
t3table). - class FilterResolver extends TreeNodeResolver[Filter, LogicalPlan] with ResolvesNameByHiddenOutput with ValidatesFilter
Resolves Filter node and its condition.
- class FunctionResolver extends TreeNodeResolver[UnresolvedFunction, Expression] with ProducesUnresolvedSubtree with CoercesExpressionTypes
A resolver for UnresolvedFunctions that resolves functions to concrete Expressions.
A resolver for UnresolvedFunctions that resolves functions to concrete Expressions. It resolves the children of the function first by calling ExpressionResolver.resolve on them if they are not UnresolvedStars. If the children are UnresolvedStars, it resolves them using ExpressionResolver.resolveStar. Examples are following:
- Function doesn't contain any UnresolvedStar:
SELECT ARRAY(col1) FROM VALUES (1);it is resolved only using ExpressionResolver.resolve.
- Function contains UnresolvedStar:
SELECT ARRAY(*) FROM VALUES (1);it is resolved using ExpressionResolver.resolveStar.
After resolving the function with FunctionResolution.resolveFunction specific expression nodes require further resolution. See resolve for more details.
Finally apply type coercion to the result of previous step and in case that the resulting expression is TimeZoneAwareExpression, apply timezone.
- class GroupingAndAggregateExpressionsExtractor extends AnyRef
Used to extract aggregate and grouping expressions in the context of underlying Aggregate operator or collecting aggregate expressions based on the provided expression.
- class HavingResolver extends TreeNodeResolver[UnresolvedHaving, LogicalPlan] with RewritesAliasesInTopLcaProject with ResolvesNameByHiddenOutput with ValidatesFilter
Resolves UnresolvedHaving node and its condition.
- class HybridAnalyzer extends SQLConfHelper
The HybridAnalyzer routes the unresolved logical plan between the legacy Analyzer and a single-pass Analyzer when the query that we are processing is being run from unit tests depending on the testing flags set and the structure of this unresolved logical plan:
The HybridAnalyzer routes the unresolved logical plan between the legacy Analyzer and a single-pass Analyzer when the query that we are processing is being run from unit tests depending on the testing flags set and the structure of this unresolved logical plan:
- If the "spark.sql.analyzer.singlePassResolver.enabled" is "true", the HybridAnalyzer will unconditionally run the single-pass Analyzer, which would usually result in some unexpected behavior and failures. This flag is used only for development.
- If the "spark.sql.analyzer.singlePassResolver.dualRunEnabled" is "true", the
HybridAnalyzer will invoke the legacy analyzer and optionally _also_ the fixed-point
one depending on the structure of the unresolved plan. This decision is based on which
features are supported by the single-pass Analyzer, and the checking is implemented in the
ResolverGuard. It's also determined if the query should be run in dual run mode by the
SQLConf.ANALYZER_DUAL_RUN_SAMPLE_RATE flag value.
If SQLConf.ANALYZER_LOG_ERRORS_INSTEAD_OF_THROWING_IN_DUAL_RUNS is enabled we tag the
query with appropriate tag. After that we validate the results using the following logic:
- If the fixed-point Analyzer fails and the single-pass one succeeds, we throw an appropriate exception (please check the QueryCompilationErrors.fixedPointFailedSinglePassSucceeded method). If SQLConf.ANALYZER_DUAL_RUN_LOGGING is enabled, we tag the query, log the message and throw the exception from the fixed-point Analyzer.
- If both the fixed-point and the single-pass Analyzers failed, we throw the exception from the fixed-point Analyzer.
- If the single-pass Analyzer failed, we throw an exception from its failure. If SQLConf.ANALYZER_DUAL_RUN_LOGGING is enabled, we tag the query, log the message and return the resolved plan from the fixed-point Analyzer.
- If both the fixed-point and the single-pass Analyzers succeeded, we compare the logical plans and output schemas, and return the resolved plan from the fixed-point Analyzer. If SQLConf.ANALYZER_DUAL_RUN_LOGGING is enabled we also tag the query.
- Otherwise we run the legacy analyzer.
- class IdentifierAndCteSubstitutor extends AnyRef
The IdentifierAndCteSubstitutor is responsible for substituting the IDENTIFIERs (not yet implemented) and CTE references in the unresolved logical plan before the actual resolution starts (specifically before metadata resolution).
The IdentifierAndCteSubstitutor is responsible for substituting the IDENTIFIERs (not yet implemented) and CTE references in the unresolved logical plan before the actual resolution starts (specifically before metadata resolution). This is important for SQL features like WITH (that could confuse MetadataResolver with extra UnresolvedRelations) or IDENTIFIER (that "hides" the actual UnresolvedRelations).
We only recurse into the plan if IdentifierAndCteSubstitutor.NODES_OF_INTEREST are present. This is done so that IdentifierAndCteSubstitutor is fast and not invasive.
- class JoinResolver extends TreeNodeResolver[Join, LogicalPlan]
Resolves Join operator by resolving its left and right children and its join condition.
Resolves Join operator by resolving its left and right children and its join condition. If the unresolved join is NaturalJoin or UsingJoin, the resulting operator will be Project, otherwise it will be Join.
- class LateralColumnAliasProhibitedRegistry extends LateralColumnAliasRegistry
Dummy implementation of LateralColumnAliasRegistry used when SQLConf.LATERAL_COLUMN_ALIAS_IMPLICIT_ENABLED is disabled.
Dummy implementation of LateralColumnAliasRegistry used when SQLConf.LATERAL_COLUMN_ALIAS_IMPLICIT_ENABLED is disabled. Getter methods throw an exception as they should never be called on a dummy implementation. Non-getter methods must remain idempotent.
- abstract class LateralColumnAliasRegistry extends AnyRef
Base class for lateral column alias registry.
Base class for lateral column alias registry. This class is extended by 2 implementations:
- LateralColumnAliasRegistryImpl - When SQLConf.LATERAL_COLUMN_ALIAS_IMPLICIT_ENABLED is enabled, this class implements logic for LCA resolution. 2. LateralColumnAliasProhibitedRegistry - Dummy class whose methods throw exceptions when LCA resolution is disabled by SQLConf.LATERAL_COLUMN_ALIAS_IMPLICIT_ENABLED.
- class LateralColumnAliasRegistryImpl extends LateralColumnAliasRegistry
LateralColumnAliasRegistryImpl is a utility class that contains structures required for lateral column alias resolution.
LateralColumnAliasRegistryImpl is a utility class that contains structures required for lateral column alias resolution. Here we store:
- currentAttributeDependencyLevelStack - Current attribute dependency level in the scope. Dependency level is defined as a maximum dependency in that attribute's expression tree. For example, in a query like:
SELECT a, b, a + b AS c, a + c AS d
Dependency levels will be as follows: level 0: a, b level 1: c level 2: d
We add a new entry to the stack for each new Alias resolution. This is needed because we can have nesting Aliases in the plan, that do not belong to the same LCA scope. For example, in the following query:
SELECT STRUCT('alpha' AS A, 'beta' AS B) ST
ST, A and B would be aliases in the same expression tree, but they do not belong in the same LCA scope.
- availableAttributes - All attributes that can be laterally referenced. This map is indexed by name, but contains a list of attributes with the same name. This is because it is possible to have multiple attributes with the same name in the scope, but they can't be laterally referenced. Handling ambiguous references is done in the getAttribute method. For the following query:
SELECT 0 AS a, 1 AS b, 2 AS c, b AS d, a AS e, d AS f, a AS g, g AS h, h AS i
availableAttributes will be: {a, b, c, d, e, f, g, h, i}
- referencedAliases - Aliases that have been laterally referenced. For the given query example, referencedAliases will be: {a, b, d, g, h}
- aliasDependencyLevels - Dependency levels of all aliases, indexed by dependency level. For the given query example, dependency levels will be as follows:
level 0: a, b, c level 1: d, e, g level 2: f, h level 3: i
- class LateralColumnAliasResolver extends QueryErrorsBase
Handles resolution of lateral column references in Project and Aggregate operators.
- class LimitLikeExpressionValidator extends QueryErrorsBase
The LimitLikeExpressionValidator validates LocalLimit, GlobalLimit, Offset or Tail integer expressions.
- type LogicalPlanResolver = TreeNodeResolver[LogicalPlan, LogicalPlan]
- class MetadataResolver extends SQLConfHelper with RelationMetadataProvider with DelegatesResolutionToExtensions
The MetadataResolver performs relation metadata resolution based on the unresolved plan at the start of the analysis phase.
The MetadataResolver performs relation metadata resolution based on the unresolved plan at the start of the analysis phase. Usually it does RPC calls to some table catalog and to table metadata itself.
RelationsWithResolvedMetadata is a map from relation ID to the relations with resolved metadata. It's produced by resolve and is used later in Resolver to replace UnresolvedRelations.
This object is one-shot per SQL query or DataFrame program resolution.
- class NameScope extends SQLConfHelper
The NameScope is used to control the resolution of names (table, column, alias identifiers).
The NameScope is used to control the resolution of names (table, column, alias identifiers). It's a part of the Resolver's state, and is used to manage the output of SQL query/DataFrame program operators.
The NameScope output is immutable. If it's necessary to update the output, NameScopeStack methods are used (overwriteCurrent or pushScope). The NameScope is always used through the NameScopeStack.
The resolution of identifiers is case-insensitive.
Name resolution priority is as follows:
- Resolution of local references:
- column reference
- parameterless function reference
- struct field or map key reference 2. Resolution of lateral column aliases (if enabled). 3. In the context of Aggregate: resolution of names in grouping expressions list referencing aliases in aggregate expressions.
Following examples showcase the priority of name resolution:
SELECT 1 AS col1, col1 FROM VALUES (2)
Because column resolution has a higher priority than LCA resolution, the result will be [1, 2] and not [1, 1].
CREATE TABLE t AS SELECT col1 as current_date FROM VALUES (2); SELECT 1 AS current_timestamp, current_timestamp, current_date FROM foo;
Result of the previous SELECT will be: [1, 2025-02-13T07:55:26.206+00:00, 2]. As can be seen, because of resolution precedence, current_date is resolved as a table column, but current_timestamp is resolved as a function without parenthesis instead of a lateral column reference.
Approximate tree of NameScope manipulations is shown in the following example:
CREATE TABLE IF NOT EXISTS t1 (col1 INT, col2 INT, col3 STRING); SELECT col1, col2 as alias1 FROM (SELECT * FROM VALUES (1, 2)) UNION (SELECT t2.col1, t2.col2 FROM (SELECT col1, col2 FROM t1) AS t2) ;
->
unionAttributes = pushScope { lhsOutput = pushScope { expandedStar = pushScope { scopes.overwriteCurrent(localRelation.output) scope.expandStar(star) } scopes.overwriteCurrent(expandedStar) scope.output } rhsOutput = pushScope { subqueryAttributes = pushScope { scopes.overwriteCurrent(t1.output) scopes.overwriteCurrent(prependQualifier(scope.output, "t2")) [scope.matchMultiPartName("t2", "col1"), scope.matchMultiPartName("t2", "col2")] } scopes.overwriteCurrent(subqueryAttributes) scope.output } scopes.overwriteCurrent(coerce(lhsOutput, rhsOutput)) [scope.matchMultiPartName("col1"), alias(scope.matchMultiPartName("col2"), "alias1")] } scopes.overwriteCurrent(unionAttributes) - Resolution of local references:
- class NameScopeStack extends SQLConfHelper
The NameScopeStack is a stack of NameScopes managed by the Resolver.
The NameScopeStack is a stack of NameScopes managed by the Resolver. Usually the current scope is used for name resolution, but in case of correlated subqueries we can lookup names in the parent scopes. Low-level scope creation is managed internally, and only high-level api like pushScope and popScope is available to the resolvers. Freshly-created NameScopeStack contains an empty root NameScope, which in the context of Resolver corresponds to the query output.
- case class NameTarget(candidates: Seq[Expression], aliasName: Option[String] = None, aliasMetadata: Option[Metadata] = None, lateralAttributeReference: Option[Attribute] = None, output: Seq[Attribute] = Seq.empty, isOuterReference: Boolean = false) extends Product with Serializable
NameTarget is a result of a multipart name resolution of the NameScope.resolveMultipartName.
NameTarget is a result of a multipart name resolution of the NameScope.resolveMultipartName.
Attribute resolution:
-- [[NameTarget]] with a single candidate `col1`. `aliasName` is be `None` in this case because -- the name is not a field/value/item of some recursive type. SELECT col1 FROM VALUES (1);
Attribute resolution ambiguity:
-- [[NameTarget]] with candidates `col1`, `col1`. [[pickCandidate]] will throw -- `AMBIGUOUS_REFERENCE`. SELECT col1 FROM VALUES (1) t1, VALUES (2) t2;
Struct field resolution:
-- [[NameTarget]] with a single candidate `GetStructField(col1, "field1")`. `aliasName` is -- `Some("col1")`, since here we extract a field of a struct. SELECT col1.field1 FROM VALUES (named_struct('field1', 1), 3);
- candidates
A list of candidates that are possible matches for a given name.
- aliasName
If the candidates size is 1 and it's type is ExtractValue (which means that it's a field/value/item from a recursive type), then the
aliasNameshould be the name with which the candidate needs to be aliased. Otherwise,aliasNameisNone.- aliasMetadata
If the candidates were created out of expressions referenced by group by alias, store the metadata of the alias. Otherwise,
aliasMetadataisNone.- lateralAttributeReference
If the candidate is laterally referencing another column this field is populated with that column's attribute.
- output
output of a NameScope that produced this NameTarget. Used to provide suggestions for thrown errors.
- isOuterReference
A flag indicating that this NameTarget resolves to an outer reference.
- sealed trait OrdinalReplacementExpressions extends AnyRef
Common trait used to store expression that are candidates to replace ordinals in grouping or sorting expressions.
- case class OrdinalReplacementGroupingExpressions(expressions: IndexedSeq[NamedExpression], hasStar: Boolean, expressionIndexesWithAggregateFunctions: HashSet[Int]) extends OrdinalReplacementExpressions with Product with Serializable
Wrapper object used to store expression that are candidates to replace ordinals in grouping expressions.
- case class OrdinalReplacementSortOrderExpressions(expressions: IndexedSeq[NamedExpression], unresolvedSort: Sort) extends OrdinalReplacementExpressions with Product with Serializable
Wrapper object used to store expression that are candidates to replace ordinals in sorting expressions.
- class OrdinalResolver extends TreeNodeResolver[UnresolvedOrdinal, Expression]
Resolves UnresolvedOrdinal to an expression from aggregate/project list, if possible.
- class PlanLogger extends Logging
PlanLogger is used by the Resolver to log intermediate resolution results.
- class PlanRewriter extends AnyRef
Utility wrapper on top of RuleExecutor, used to apply post-resolution rules on single-pass resolution result.
Utility wrapper on top of RuleExecutor, used to apply post-resolution rules on single-pass resolution result. SinglePassRewriter transforms the plan and the subqueries inside.
- trait ProducesUnresolvedSubtree extends ResolvesExpressionChildren
A mixin trait for expression resolvers that as part of their resolution, replace single node with a subtree of nodes.
A mixin trait for expression resolvers that as part of their resolution, replace single node with a subtree of nodes. This step is necessary because the underlying legacy code that is being called produces partially-unresolved subtrees. In order to resolve the subtree a callback resolver is called recursively. This callback must ensure that no node is resolved twice in order to not break the single-pass invariant. This is done by tagging the limits of this traversal with ResolverTag.SINGLE_PASS_SUBTREE_BOUNDARY tag. This tag is applied to the original expression's children, which are guaranteed to be resolved at the time of given expression's resolution. When callback resolver encounters the node that is tagged, it should return identity instead of trying to resolve it.
- class ProhibitedResolver extends LogicalPlanResolver
This is a dummy LogicalPlanResolver whose resolve is not implemented and throws SparkException.
This is a dummy LogicalPlanResolver whose resolve is not implemented and throws SparkException.
It's used by the MetadataResolver to pass it as an argument to tryDelegateResolutionToExtensions, because unresolved subtree resolution doesn't make sense during metadata resolution traversal.
- class ProjectResolver extends TreeNodeResolver[Project, LogicalPlan]
Resolves initially unresolved Project operator to either a resolved Project or Aggregate node, based on whether there are aggregate expressions in the project list.
Resolves initially unresolved Project operator to either a resolved Project or Aggregate node, based on whether there are aggregate expressions in the project list. When LateralColumnAlias resolution is enabled, replaces the output operator with an appropriate operator structure using information from the scope. Detailed explanation can be found in buildProjectWithResolvedLCAs method.
- case class RelationId(multipartIdentifier: Seq[String], options: CaseInsensitiveStringMap = CaseInsensitiveStringMap.empty, isStreaming: Boolean = false) extends Product with Serializable
The RelationId is a unique identifier for a relation.
The RelationId is a unique identifier for a relation. It is used to lookup the relations which were processed by the MetadataResolver to substitute the unresolved relations in single pass during the analysis phase.
- trait RelationMetadataProvider extends LookupCatalog
RelationMetadataProvider provides relations with resolved metadata based on the corresponding UnresolvedRelations.
RelationMetadataProvider provides relations with resolved metadata based on the corresponding UnresolvedRelations. It is used by Resolver to replace UnresolvedRelation with a specific LogicalPlan with resolved metadata, e.g. with UnresolvedCatalogRelation or View.
- class ResolutionCheckRunner extends SQLConfHelper
The ResolutionCheckRunner is used to run
resolutionCheckson the logical plan.The ResolutionCheckRunner is used to run
resolutionCheckson the logical plan.Important note: these checks are not always idempotent, and sometimes perform heavy network operations.
- class ResolutionValidator extends AnyRef
The ResolutionValidator performs the validation work after the logical plan tree is resolved by the Resolver.
The ResolutionValidator performs the validation work after the logical plan tree is resolved by the Resolver. Each
resolve*method in the Resolver must have itsvalidate*counterpart in the ResolutionValidator. The validation code asserts the conditions that must never be false no matter which SQL query or DataFrame program was provided. The validation approach is single-pass, post-order, complementary to the resolution process. - case class ResolvedAggregateExpressions(expressions: Seq[NamedExpression], resolvedExpressionsWithoutAggregates: Seq[NamedExpression], hasAttributeOutsideOfAggregateExpressions: Boolean, hasStar: Boolean, expressionIndexesWithAggregateFunctions: HashSet[Int], hasLateralColumnAlias: Boolean) extends Product with Serializable
ResolvedAggregateExpressions is used by the ExpressionResolver.resolveAggregateExpressions to return resolution results.
ResolvedAggregateExpressions is used by the ExpressionResolver.resolveAggregateExpressions to return resolution results.
- expressions: The resolved expressions. They are resolved using the
resolveExpressionTreeInOperator. - resolvedExpressionsWithoutAggregates: List of resolved aggregate expressions that don't have AggregateExpressions in their subtrees.
- hasAttributeOutsideOfAggregateExpressions: True if
expressionslist contains any attributes that are not under an AggregateExpression. - hasStar: True if there is a star (
*) in aggregate expressions list - expressionIndexesWithAggregateFunctions: Indices of expressions in aggregate expressions list that have aggregate functions in their subtrees.
- hasLateralColumnAlias: True if there is a lateral column reference in the aggregate expressions list.
- expressions: The resolved expressions. They are resolved using the
- case class ResolvedProjectList(expressions: Seq[NamedExpression], hasAggregateExpressions: Boolean, hasLateralColumnAlias: Boolean, aggregateListAliases: Seq[Alias], baseAggregate: Option[Aggregate] = None) extends Product with Serializable
Structure used to return results of the resolved project list.
Structure used to return results of the resolved project list.
- expressions: The resolved expressions. It is resolved using the
resolveExpressionTreeInOperator. - hasAggregateExpressions: True if the resolved project list contains any aggregate expressions.
- hasLateralColumnAlias: True if the resolved project list contains any lateral column aliases.
- aggregateListAliases: List of aliases in aggregate list if there are aggregate expressions in the Project.
- baseAggregate: Base Aggregate node constructed by LateralColumnAliasResolver while resolving lateral column references in Aggregate.
- expressions: The resolved expressions. It is resolved using the
- case class ResolvedSubqueryExpressionPlan(plan: LogicalPlan, output: Seq[Attribute], outerExpressions: Seq[Expression]) extends Product with Serializable
The result of SubqueryExpression.plan resolution.
The result of SubqueryExpression.plan resolution. This is used internally in SubqueryExpressionResolver.
- plan
The resolved plan of the subquery.
- output
Plan output. We don't use LogicalPlan.output in the single-pass Analyzer, because this method is often recursive.
- outerExpressions
The outer expressions that are references in the plan. OuterReference wrapper is stripped away. These can be either actual leaf AttributeReferences or AggregateExpressions with outer references inside.
- class Resolver extends LogicalPlanResolver with DelegatesResolutionToExtensions
The Resolver implements a single-pass bottom-up analysis algorithm in the Catalyst.
The Resolver implements a single-pass bottom-up analysis algorithm in the Catalyst.
The functions here generally traverse the LogicalPlan nodes recursively, constructing and returning the resolved LogicalPlan nodes bottom-up. This is the primary entry point for implementing SQL and DataFrame plan analysis, wherein the resolve method accepts a fully unresolved LogicalPlan and returns a fully resolved LogicalPlan in response with all data types and attribute reference ID assigned for valid requests. This resolver also takes responsibility to detect any errors in the initial SQL query or DataFrame and return appropriate error messages including precise parse locations wherever possible.
The Resolver is a one-shot object per each SQL/DataFrame logical plan, the calling code must re-create it for every new analysis run.
- trait ResolverExtension extends AnyRef
The ResolverExtension is a main interface for single-pass analysis extensions in Catalyst.
The ResolverExtension is a main interface for single-pass analysis extensions in Catalyst. External code that needs specific node types to be resolved has to implement this trait and inject the implementation into the Analyzer.singlePassResolverExtensions.
- class ResolverGuard extends SQLConfHelper
ResolverGuard is a class that checks if the operator that is yet to be analyzed only consists of operators and expressions that are currently supported by the single-pass analyzer.
ResolverGuard is a class that checks if the operator that is yet to be analyzed only consists of operators and expressions that are currently supported by the single-pass analyzer.
This is a one-shot object and should not be reused after apply call.
- trait ResolverMetricTracker extends AnyRef
Trait for tracking and logging timing metrics for single-pass resolver.
- class ResolverRunner extends ResolverMetricTracker with SQLConfHelper
Wrapper class for Resolver and single-pass resolution.
Wrapper class for Resolver and single-pass resolution. This class encapsulates single-pass resolution, rewriting and validation of resolved plan. The plan rewrite is necessary in order to either fully resolve the plan or stay compatible with the fixed-point analyzer.
- trait ResolvesExpressionChildren extends AnyRef
- trait ResolvesNameByHiddenOutput extends SQLConfHelper
ResolvesNameByHiddenOutput is used by resolvers for operators that are able to resolve attributes in its expression tree from hidden output or that can reference expressions not present in child's output.
ResolvesNameByHiddenOutput is used by resolvers for operators that are able to resolve attributes in its expression tree from hidden output or that can reference expressions not present in child's output. Update child operator's output list and place a Project node on top of original operator node with the original output of an operator's child.
For example, in a following query:
SELECT t1.key FROM t1 FULL OUTER JOIN t2 USING (key) WHERE t1.key NOT LIKE 'bb.%';Plan without adding missing attributes would be:
+- Project [key#1] +- Filter NOT key#1 LIKE bb.% +- Project [coalesce(key#1, key#2) AS key#3, __key#1__, __key#2__] +- Join FullOuter, (key#1 = key#2) :- SubqueryAlias t1 : +- Relation t1[key#1] +- SubqueryAlias t2 +- Relation t2[key#2]
NOTE: #key1 and key#2 at the end of inner Project are metadata columns from the full outer join. Even though they are present in the Project in single-pass, fixed-point adds these columns after resolving missing input, so duplication of some columns is possible. In order to stay fully compatible between single-pass and fixed-point, we add both missing attributes and these metadata columns. We mimic fixed-point behavior by putting metadata columns in instead of NameScope.output.
In the plan above, Filter requires key#1 in its condition, but key#1 is not available in the below Project's output, even though key#1 is available in Join's hidden output. Because of that, we need to place key#1 in the project list, after original project list expressions, but before metadata columns (to remain compatible with fixed-point). In order to preserve initial output of Filter, we place a Project node on top of this Filter, whose project list is the original output of the Project below Filter (in this case - key#3 and metadata columns key#1 and key#2).
Therefore, the plan becomes:
+- Project [key#1] +- Project [key#3, key#1, key#2] +- Filter NOT key#1 LIKE bb.% +- Project [coalesce(key#1, key#2) AS key#3, key#1, key#1, key#2] +- Join FullOuter, (key#1 = key#2) :- SubqueryAlias t1 : +- Relation t1[key#1] +- SubqueryAlias t2 +- Relation t2[key#2]
Query below exhibits similar behavior when Sort operator resolves an attribute using hidden output:
SELECT col1 FROM VALUES (1, 2) ORDER BY col2;
Unresolved plan would be:
Sort [col2 ASC NULLS FIRST], true +- Project [col1] +- LocalRelation [col1, col2]As it can be seen, attribute
col2used in Sort can't be resolved using the Project output (which is [col1]), so it has to be resolved using the hidden output (which is propagated from LocalRelation and is [col1,col2]). As it's been shown in the previous example,col2has to be added to Project list and a Project with original output of the Project below Sort is added as a top node. Because of that, analyzed plan is:Project [col1] +- Sort [col2 ASC NULLS FIRST], true +- Project [col1, col2] +- LocalRelation [col1, col2]Another example is when Sort order expression is an AggregateExpression which is not present in the Aggregate.aggregateExpressions:
SELECT col1 FROM VALUES (1) GROUP BY col1 ORDER BY sum(col1);In this example
sum(col1)should be added to child's output and a Project node should be added on top of the Sort node to preserve the original output of the Aggregate node:Project [col1] +- Sort [sum(col1)#... ASC NULLS FIRST], true +- Aggregate [col1], [col1, sum(col1) AS sum(col1)#...] +- LocalRelation [col1]
In case of Dataframe programs we can have multiple Sort operators nested inside each other. For example:
df.select("col1").orderBy("col2").orderBy("col1").orderBy("col2")
Unresolved plan would be:
Sort [col2 ASC NULLS FIRST], true +- Sort [col1 ASC NULLS FIRST], true +- Project [col1] +- Sort [col2 ASC NULLS FIRST], true +- Project [col1, col2] +- Project [col1, col2, col3] +- LocalRelation [col1, col2, col3]
As it can be seen,
col2(Sort order expression) needs to be resolved using the hidden output. Because of that it must be added to all the Projects and Aggregates below the Sort operator. Resolved plan would be:Project [col1] +- Sort [col2 ASC NULLS FIRST], true +- Sort [col1 ASC NULLS FIRST], true +- Project [col1, col2] +- Sort [col2 ASC NULLS FIRST], true +- Project [col1, col2] +- Project [col1, col2, col3] +- LocalRelation [col1, col2, col3]
In the plan you can see that
col2is added to the lower Project.projectList. - trait RewritesAliasesInTopLcaProject extends AnyRef
During LCA resolution some aliases may be rewritten as new aliases with new ExprIds.
During LCA resolution some aliases may be rewritten as new aliases with new ExprIds. This trait handles remapping of old aliases to new ones, when these attributes appear in SortOrder expressions and Having conditions.
- class SemanticComparator extends AnyRef
SemanticComparator is a tool to compare expressions semantically to a predefined sequence of
targetExpressions.SemanticComparator is a tool to compare expressions semantically to a predefined sequence of
targetExpressions. Semantic comparison is based on QueryPlan.canonicalized - for example,col1 + 1 + col2is semantically equal to1 + col2 + col1. To speed up slow tree traversals and expression node field comparisons, we cache the semantic hashes (which is simply a hash of a canonicalized subtree) and use them for O(1) indexing. If the hashes don't match, we perform an early return. Otherwise, we invoke the heavy Expression.semanticEquals method to make sure that expression trees are indeed identical. - class SemiStructuredExtractResolver extends TreeNodeResolver[SemiStructuredExtract, Expression] with ResolvesExpressionChildren with CoercesExpressionTypes
Resolver for SemiStructuredExtract.
Resolver for SemiStructuredExtract. Resolves SemiStructuredExtract by resolving its children, replacing it with the proper semi-structured field extraction method and applying type coercion to the result.
- class SetOperationLikeResolver extends TreeNodeResolver[LogicalPlan, LogicalPlan]
The SetOperationLikeResolver performs Union, Intersect or Except operator resolution.
The SetOperationLikeResolver performs Union, Intersect or Except operator resolution. These operators have 2+ children. Resolution involves checking and normalizing child output attributes (data types and nullability).
- class SortResolver extends TreeNodeResolver[Sort, LogicalPlan] with RewritesAliasesInTopLcaProject with ResolvesNameByHiddenOutput
Resolves a Sort by resolving its child and order expressions.
- class SubqueryExpressionResolver extends CoercesExpressionTypes
SubqueryExpressionResolver resolves specific SubqueryExpressions, such as ScalarSubquery, ListQuery, and Exists.
- class SubqueryRegistry extends AnyRef
The SubqueryRegistry manages the stack of SubqueryScopes during the resolution of the whole SQL query.
The SubqueryRegistry manages the stack of SubqueryScopes during the resolution of the whole SQL query. Every new SubqueryScope has its own isolated scope.
- class SubqueryScope extends AnyRef
The SubqueryScope is managed through the whole resolution process of a given SubqueryExpression plan.
The SubqueryScope is managed through the whole resolution process of a given SubqueryExpression plan.
The reason why we need this scope is that AggregateExpressions with OuterReferences are handled in a special way. Consider this query:
-- t1.col2 is an outer reference SELECT col1 FROM VALUES (1, 2) t1 GROUP BY col1 HAVING ( SELECT * FROM VALUES (1, 2) t2 WHERE t2.col2 == MAX(t1.col2) )
During the Exists resolution inside the HAVING clause we encounter "t1.col2" name, which is resolved to an OuterReference. There's an AggregateExpression on top of it. This whole expression is not local to the subquery, and thus it belongs to an outer Aggregate operator below the
HAVINGclause. We need top pull it up outside of the subquery, and insert it in the Aggregate operator. So the resolution order is as follows:- Resolve "t1.col2" to an OuterReference in ExpressionResolver.resolveAttribute;
- Resolve the AggregateExpression in AggregateExpressionResolver.resolve;
- Detect an outer reference below the aggregate expression, cut the whole subtree with outer references stripped away, alias it and insert it in this SubqueryScope;
- Replace the aggregate expression with an OuterReference to the AttributeReference from that artificial Alias;
- When the resolution of the SubqueryExpression is finished, SubqueryRegistry.popScope merges the lower scope to the upper one, and all the outer aggregate expression references are appended to the common lowerAliasedOuterAggregateExpressions list.
- Finally, the resolution of the
HAVINGclause can insert the missing aggregate expression into the lower Aggregate operator. During this process we must call ExpressionIdAssigner.mapExpression on the new alias, because this auto-generated alias is new to the query plan, so that ExpressionIdAssigner remembers it.
Notes:
- Spark only supports outer aggregates in the subqueries inside
HAVING; - The subtree under a given AggregateExpression can be arbitrary, but must contain either local or outer references, the mixed set is disallowed.
- We can have several subquery expressions in HAVING clause, that's why we append outer aggregate expressions from lower scopes in mergeChildScope.
- class TimezoneAwareExpressionResolver extends TreeNodeResolver[TimeZoneAwareExpression, Expression] with ResolvesExpressionChildren with CoercesExpressionTypes
Resolves TimeZoneAwareExpressions by applying the session's local timezone.
Resolves TimeZoneAwareExpressions by applying the session's local timezone.
This class is responsible for resolving TimeZoneAwareExpressions by first resolving their children and then applying the session's local timezone. Additionally, ensures that any tags from the original expression are preserved during the resolution process.
- trait TreeNodeResolver[UnresolvedNode <: TreeNode[_], ResolvedNode <: TreeNode[_]] extends SQLConfHelper with QueryErrorsBase
Base class for TreeNode resolvers.
Base class for TreeNode resolvers. All resolvers should extend this class with specific UnresolvedNode and ResolvedNode types.
- case class UnresolvedCteRelationRef(name: String) extends LogicalPlan with UnresolvedLeafNode with NamedRelation with Product with Serializable
A reference to a CTE definition in the form of an unresolved relation.
A reference to a CTE definition in the form of an unresolved relation. This node is introduced by IdentifierAndCteSubstitutor to replace the CTE reference to avoid ineffective catalog RPC lookups in MetadataResolver.
- trait ValidatesFilter extends QueryErrorsBase
ValidatesFilter is used by resolvers to validate resolved Filter operator.
- case class ViewResolutionContext(nestedViewDepth: Int, maxNestedViewDepth: Int, collation: Option[String] = None, catalogAndNamespace: Option[Seq[String]] = None) extends Product with Serializable
The ViewResolutionContext consists of data, which is specific to the specific view plan resolution.
The ViewResolutionContext consists of data, which is specific to the specific view plan resolution. This data is also propagated to the subviews.
- nestedViewDepth
Current nested view depth. Cannot exceed the
maxNestedViewDepth.- maxNestedViewDepth
Maximum allowed nested view depth. Configured in the upper context based on SQLConf.MAX_NESTED_VIEW_DEPTH.
- collation
View's default collation if explicitly set.
- catalogAndNamespace
Catalog and camespace under which the View was created.
- class ViewResolver extends TreeNodeResolver[View, View]
The ViewResolver resolves view plans that were already reconstructed by SessionCatalog from the view text and view metadata (schema, configs).
Value Members
- object AnalyzerBridgeState extends Serializable
- object CoercesExpressionTypes
- object CteRegistry
- object DefaultCollationTypeCoercion
This type coercion object is only used in the single-pass analyzer and is not part of the TypeCoercion rules used in the fixed-point analyzer.
This type coercion object is only used in the single-pass analyzer and is not part of the TypeCoercion rules used in the fixed-point analyzer.
When the database object (e.g. View) has a custom default collation and a resolving expression's dataType is the companion object StringType, we need to treat the dataType as a collated StringType. To do this, we wrap the expression with a cast to a StringType with the default collation or change the dataType to StringType with the default collation. Note that when the dataType is equal to the companion object StringType, but isn't the same by reference, we shouldn't change the dataType, since that means the user explicitly specified UTF8_BINARY collation.
- object ExplicitlyUnsupportedResolverFeature extends Serializable
This object contains all the metadata on explicitly unsupported resolver features.
- object ExpressionIdAssigner
- object ExpressionResolutionContext
- object HybridAnalyzer
- object IdentifierAndCteSubstitutor
- object OperatorWithUncomparableTypeValidator
OperatorWithUncomparableTypeValidator performs the validation of a logical plan to ensure that it (if it is Distinct or SetOperation) does not contain any uncomparable types: VariantType, MapType, GeometryType or GeographyType.
- object OutputType extends Enumeration
OutputType represents different types of output used during multipart name resolution in the NameScope.
- object PruneMetadataColumns extends Rule[LogicalPlan]
This is a special rule for single-pass resolver that performs a single downwards traversal in order to prune away unnecessary metadata columns.
This is a special rule for single-pass resolver that performs a single downwards traversal in order to prune away unnecessary metadata columns. This is necessary because fixed-point is looking into the operator tree in order to determine whether it is necessary to add metadata columns. In single-pass, we always add metadata columns during the main traversal and this rule performs the cleanup of those columns that are unnecessary. Important thing to note here is that by "unnecessary" columns we are not referring to the ones that are not needed in upper operators for correct result, but to the columns that have been added by single-pass resolver but are not present in the fixed-point plan.
- object Resolver
- object ResolverGuard
- object ResolverMetricTracker
- object ResolverTag
Object used to store single-pass resolver related tags.
- object TimezoneAwareExpressionResolver
- object UnsupportedExpressionInOperatorValidation