class NameScope extends SQLConfHelper
The NameScope is used to control the resolution of names (table, column, alias identifiers). It's a part of the Resolver's state, and is used to manage the output of SQL query/DataFrame program operators.
The NameScope output is immutable. If it's necessary to update the output, NameScopeStack methods are used (overwriteCurrent or pushScope). The NameScope is always used through the NameScopeStack.
The resolution of identifiers is case-insensitive.
Name resolution priority is as follows:
- Resolution of local references:
- column reference
- parameterless function reference
- struct field or map key reference 2. Resolution of lateral column aliases (if enabled). 3. In the context of Aggregate: resolution of names in grouping expressions list referencing aliases in aggregate expressions.
Following examples showcase the priority of name resolution:
SELECT 1 AS col1, col1 FROM VALUES (2)
Because column resolution has a higher priority than LCA resolution, the result will be [1, 2] and not [1, 1].
CREATE TABLE t AS SELECT col1 as current_date FROM VALUES (2); SELECT 1 AS current_timestamp, current_timestamp, current_date FROM foo;
Result of the previous SELECT will be: [1, 2025-02-13T07:55:26.206+00:00, 2]. As can be seen, because of resolution precedence, current_date is resolved as a table column, but current_timestamp is resolved as a function without parenthesis instead of a lateral column reference.
Approximate tree of NameScope manipulations is shown in the following example:
CREATE TABLE IF NOT EXISTS t1 (col1 INT, col2 INT, col3 STRING); SELECT col1, col2 as alias1 FROM (SELECT * FROM VALUES (1, 2)) UNION (SELECT t2.col1, t2.col2 FROM (SELECT col1, col2 FROM t1) AS t2) ;
->
unionAttributes = pushScope {
lhsOutput = pushScope {
expandedStar = pushScope {
scopes.overwriteCurrent(localRelation.output)
scope.expandStar(star)
}
scopes.overwriteCurrent(expandedStar)
scope.output
}
rhsOutput = pushScope {
subqueryAttributes = pushScope {
scopes.overwriteCurrent(t1.output)
scopes.overwriteCurrent(prependQualifier(scope.output, "t2"))
[scope.matchMultiPartName("t2", "col1"), scope.matchMultiPartName("t2", "col2")]
}
scopes.overwriteCurrent(subqueryAttributes)
scope.output
}
scopes.overwriteCurrent(coerce(lhsOutput, rhsOutput))
[scope.matchMultiPartName("col1"), alias(scope.matchMultiPartName("col2"), "alias1")]
}
scopes.overwriteCurrent(unionAttributes)- Alphabetic
- By Inheritance
- NameScope
- SQLConfHelper
- AnyRef
- Any
- Hide All
- Show All
- Public
- Protected
Instance Constructors
- new NameScope(output: Seq[Attribute] = Seq.empty, hiddenOutput: Seq[Attribute] = Seq.empty, isSubqueryRoot: Boolean = false, availableAliases: HashSet[ExprId] = new HashSet[ExprId], aggregateListAliases: Seq[Alias] = Seq.empty, baseAggregate: Option[Aggregate] = None, planLogger: PlanLogger = new PlanLogger)
- output
These are the attributes visible for lookups in the current scope. These may be:
- Transformed outputs of lower scopes (e.g. type-coerced outputs of Union's children).
- Output of a current operator that is being resolved (leaf nodes like Relations).
- hiddenOutput
Attributes that are not directly visible in the scope, but available for lookup in case the resolved attribute is not found in
output.- isSubqueryRoot
Indicates that the current scope is a root of a subquery. This is used by NameScopeStack.resolveMultipartName to detect the nearest outer scope.
- availableAliases
User specified aliases that are present in this NameScope.
- aggregateListAliases
List of aliases that are present in the Aggregate corresponding to this NameScope. If the Aggregate has lateral column references, this list contains both the aliases from Aggregate as well as all aliases from artificially inserted Project nodes.
- baseAggregate
Aggregate node that is either a resolved Aggregate corresponding to this node or base Aggregate constructed when resolving lateral column references in Aggregate.
- planLogger
PlanLogger used to log name resolution events.
Value Members
- final def !=(arg0: Any): Boolean
- Definition Classes
- AnyRef → Any
- final def ##: Int
- Definition Classes
- AnyRef → Any
- final def ==(arg0: Any): Boolean
- Definition Classes
- AnyRef → Any
- def addTopAggregateExpression(aliasedAggregateExpression: Alias): Unit
Add a top level alias to the map so it can be used when resolving a grouping expression.
Add a top level alias to the map so it can be used when resolving a grouping expression. We store Aliases with a same name in a list related to it as the order of the definition of the aliases in the aggregate expressions list is important (first one should be used for resolution in group by alias case). Example:
SELECT col1 as a, col2 AS a FROM values('a', 'b') GROUP BY col2, a;
afrom the grouping expressions should be resolved ascol1(as it comes beforecol2 AS ain the aggregate expressions list). Plan looks like:Aggregate [col2#2, col1#1], [col1#1 AS a#3, col2#2 AS a#4] +- LocalRelation [col1#1, col2#2]
- val aggregateListAliases: Seq[Alias]
- final def asInstanceOf[T0]: T0
- Definition Classes
- Any
- val availableAliases: HashSet[ExprId]
- val baseAggregate: Option[Aggregate]
- def clone(): AnyRef
- Attributes
- protected[lang]
- Definition Classes
- AnyRef
- Annotations
- @throws(classOf[java.lang.CloneNotSupportedException]) @IntrinsicCandidate() @native()
- def conf: SQLConf
The active config object within the current scope.
The active config object within the current scope. See SQLConf.get for more information.
- Definition Classes
- SQLConfHelper
- final def eq(arg0: AnyRef): Boolean
- Definition Classes
- AnyRef
- def equals(arg0: AnyRef): Boolean
- Definition Classes
- AnyRef → Any
- def expandStar(unresolvedStar: Star): Seq[NamedExpression]
Expand the Star.
Expand the Star. The expected use case for this method is star expansion inside Project.
Star without a target:
-- Here the star will be expanded to [a, b, c]. SELECT * FROM VALUES (1, 2, 3) AS t(a, b, c);
Star with a multipart name target:
USE CATALOG catalog1; USE DATABASE database1; CREATE TABLE IF NOT EXISTS table1 (col1 INT, col2 INT); -- Here the star will be expanded to [col1, col2]. SELECT catalog1.database1.table1.* FROM catalog1.database1.table1;
Star with a struct target:
-- Here the star will be expanded to [field1, field2]. SELECT d.* FROM VALUES (named_struct('field1', 1, 'field2', 2)) AS t(d);
Star as an argument to a function:
-- Here the star will be expanded to [col1, col2, col3] and those would be passed as -- arguments to `concat_ws`. SELECT concat_ws('', *) AS result FROM VALUES (1, 2, 3);
Also, see Star.expandStar for more details.
- def findAttributesByName(name: String): Seq[Attribute]
Find attributes in this NameScope that match a provided one-part
name.Find attributes in this NameScope that match a provided one-part
name.This method is simpler and more lightweight than resolveMultipartName, because here we just return all the attributes matched by the one-part
name. This is only suitable for situations where name _resolution_ is not required (e.g. accessing struct fields from the lower operator's output).For example, this method is used to look up attributes to match a specific View schema. See ExpressionResolver.resolveGetViewColumnByNameAndOrdinal for more info on view column lookup.
We are relying on a simple IdentifierMap to perform that work, since we just need to match one-part name from the lower operator's output here.
- def getAttributeById(expressionId: ExprId): Option[Attribute]
Returns attribute with
expressionIdifoutputcontains it.Returns attribute with
expressionIdifoutputcontains it. This is used to preserve nullability for resolved AttributeReference. - final def getClass(): Class[_ <: AnyRef]
- Definition Classes
- AnyRef → Any
- Annotations
- @IntrinsicCandidate() @native()
- def getHiddenAttributeById(expressionId: ExprId): Option[Attribute]
Returns attribute with
expressionIdifhiddenOutputcontains it. - def getOrdinalReplacementExpressions: Option[OrdinalReplacementExpressions]
Get expressions that are candidates to be resolved by ordinal in current NameScope.
- def getOutputIds: Set[ExprId]
Return all the explicitly outputted expression IDs.
Return all the explicitly outputted expression IDs. Hidden or metadata output are not included.
- def getTopAggregateExpressionAliases: Seq[Alias]
Returns all the top-level aliases from Aggregate corresponding to this NameScope collected while resolving aggregate expressions.
Returns all the top-level aliases from Aggregate corresponding to this NameScope collected while resolving aggregate expressions. Used to resolve group by alias.
- def hashCode(): Int
- Definition Classes
- AnyRef → Any
- Annotations
- @IntrinsicCandidate() @native()
- val hiddenOutput: Seq[Attribute]
- final def isInstanceOf[T0]: Boolean
- Definition Classes
- Any
- val isSubqueryRoot: Boolean
- lazy val lcaRegistry: LateralColumnAliasRegistry
- final def ne(arg0: AnyRef): Boolean
- Definition Classes
- AnyRef
- final def notify(): Unit
- Definition Classes
- AnyRef
- Annotations
- @IntrinsicCandidate() @native()
- final def notifyAll(): Unit
- Definition Classes
- AnyRef
- Annotations
- @IntrinsicCandidate() @native()
- val output: Seq[Attribute]
- def overwrite(output: Option[Seq[Attribute]] = None, hiddenOutput: Option[Seq[Attribute]] = None, availableAliases: Option[HashSet[ExprId]] = None, aggregateListAliases: Seq[Alias] = Seq.empty, baseAggregate: Option[Aggregate] = None): NameScope
Returns new NameScope which preserves all the immutable NameScope properties but overwrites
output,hiddenOutput,availableAliases,aggregateListAliasesandbaseAggregateif provided. - def resolveMissingAttributesByHiddenOutput(referencedAttributes: HashMap[ExprId, Attribute]): Seq[Attribute]
Given referenced attributes, returns all attributes that are referenced and missing from current output, but can be found in hidden output.
- def resolveMultipartName(multipartName: Seq[String], canLaterallyReferenceColumn: Boolean = false, canReferenceAggregateExpressionAliases: Boolean = false, canResolveNameByHiddenOutput: Boolean = false, shouldPreferHiddenOutput: Boolean = false, canReferenceAggregatedAccessOnlyAttributes: Boolean = false): NameTarget
Resolve multipart name into a NameTarget.
Resolve multipart name into a NameTarget. NameTarget's
candidatesmay contain simple AttributeReferences if it's a column or alias, or ExtractValue expressions if it's a struct field, map value or array value. ThealiasNamewill optionally be set to the proposed alias name for the value extracted from a struct, map or array.Example that demonstrates those major use-cases:
CREATE TABLE IF NOT EXISTS t ( col1 INT, col2 STRUCT<field: INT>, col3 STRUCT<struct: STRUCT<field: INT>>, col4 MAP<STRING: INT>, col5 STRING ); -- For the SELECT below the top Project list will be resolved using this method like this: -- AttributeReference(col1), -- AttributeReference(a), -- GetStructField(col2, field), -- GetStructField(GetStructField(col3, struct), field), -- GetMapValue(col4, key) SELECT col1, a, col2.field, col3.struct.field, col4.key FROM (SELECT *, col5 AS a FROM t);
Since there can be several expressions that matched the same multipart name, this method may return a NameTarget with the following
candidates: - 0 values: No matched expressions - 1 value: Unique expression matched - 1+ values: Ambiguity, several expressions matchedSome examples of ambiguity:
CREATE TABLE IF NOT EXISTS t1 (c1 INT, c2 INT); CREATE TABLE IF NOT EXISTS t2 (c2 INT, c3 INT); -- Identically named columns from different tables. -- This will fail with AMBIGUOUS_REFERENCE error. SELECT c2 FROM t1, t2;CREATE TABLE IF NOT EXISTS foo (c1 INT); CREATE TABLE IF NOT EXISTS bar (foo STRUCT<c1: INT>); -- Ambiguity between a column in a table and a field in a struct. -- This will succeed, and column will win over the struct field. SELECT foo.c1 FROM foo, bar;
The candidates are deduplicated by expression ID (not by attribute name!):
CREATE TABLE IF NOT EXISTS t1 (col1 STRING); -- No ambiguity here, since we are selecting the same column (same expression ID). SELECT col1 FROM (SELECT col1, col1 FROM t);
The case of the
multipartNametakes precedence over the original name case, so the candidates will have names that are case-identical to themultipartName:CREATE TABLE IF NOT EXISTS t1 (col1 STRING); -- The output schema of this query is [COL1], despite the fact that the column is in -- lower-case. SELECT COL1 FROM t;
Name resolution can be done using the hidden output for certain operators (e.g Sort, Filter). This is indicated by
canResolveNameByHiddenOutputwhich is passed from ExpressionResolver.resolveAttribute based on the parent operator. Example:-- Project's output = [`col1`]; Project's hidden output = [`col1`, `col2`] SELECT col1 FROM VALUES(1, 2) ORDER BY col2;
When resolving Sort, SortOrder expressions should first be attempted to be resolved by table columns in project list and only then by aliases from the project list. For example, in the following query:
SELECT 1 AS col1, col1 FROM VALUES(1) ORDER BY col1
Even though there is ambiguity with the name
col1, the SortOrder expression should be resolved as a table column from the project list and not throw AMBIGUOUS_REFERENCE.The names in Aggregate.groupingExpressions can reference Aggregate.aggregateExpressions aliases.
canReferenceAggregateExpressionAliaseswill be true when we are resolving the grouping expressions. Example:{{ SELECT col1 + col2 AS a FROM VALUES (1, 2) GROUP BY a; }}}
In case we are resolving names in expression trees from HAVING or ORDER BY on top of Aggregate, we are able to resolve hidden attributes only if those are present in grouping expressions, or if the reference itself is under an AggregateExpression. In the latter case
canReferenceAggregatedAccessOnlyAttributeswill be true, and all the attributes fromhiddenOutputwill be used instead of the former case where we don't use attributes marked asaggregatedAccessOnly. Consider the following example:-- This succeeds, because `col2` is in the grouping expressions. SELECT COUNT(col1) FROM t1 GROUP BY col1, col2 ORDER BY col2; -- This fails, because `col2` is not in the grouping expressions. SELECT COUNT(col1) FROM t1 GROUP BY col1 ORDER BY col2; -- This succeeds, despite the fact that `col2` is not in the grouping expressions. -- Such references are allowed under an aggregate expression (MAX). SELECT COUNT(col1) FROM t1 GROUP BY col1 ORDER BY MAX(col2);
Spark is being smart about name resolution and prioritizes candidates from output levels that can actually be resolved, even though that output level might not be the first choice. For example, ORDER BY clause prefers attributes from SELECT list (namely, aliases) over table columns from below. However, if attributes on the SELECT level have name ambiguity or other issues, Spark will try to resolve the name using the table columns from below. Examples:
CREATE TABLE t1 (col1 INT); CREATE TABLE t2 (col1 STRUCT<field: INT>); -- Main output is ambiguous, so col1 from t1 is used for sorting. SELECT 1 AS col1, 2 AS col1 FROM t1 ORDER BY col1; -- col1 from main output does not have `field`, so struct field of col1 from t2 is used for -- sorting. SELECT 1 AS col1 FROM t2 ORDER BY col1.field;
This is achieved using candidate prioritization mechanism in pickSuitableCandidates.
We are relying on the AttributeSeq to perform name resolution, since it requires complex resolution logic involving nested field extraction and multipart name matching. See AttributeSeq.resolve for more details.
- def setOrdinalReplacementExpressions(ordinalReplacementExpressions: OrdinalReplacementExpressions): Unit
Set expressions that are candidates to be resolved by ordinal in current NameScope.
- final def synchronized[T0](arg0: => T0): T0
- Definition Classes
- AnyRef
- def toString(): String
- Definition Classes
- AnyRef → Any
- final def wait(arg0: Long, arg1: Int): Unit
- Definition Classes
- AnyRef
- Annotations
- @throws(classOf[java.lang.InterruptedException])
- final def wait(arg0: Long): Unit
- Definition Classes
- AnyRef
- Annotations
- @throws(classOf[java.lang.InterruptedException]) @native()
- final def wait(): Unit
- Definition Classes
- AnyRef
- Annotations
- @throws(classOf[java.lang.InterruptedException])
- def withSQLConf[T](pairs: (String, String)*)(f: => T): T
Sets all SQL configurations specified in
pairs, callsf, and then restores all SQL configurations.Sets all SQL configurations specified in
pairs, callsf, and then restores all SQL configurations.- Attributes
- protected
- Definition Classes
- SQLConfHelper
Deprecated Value Members
- def finalize(): Unit
- Attributes
- protected[lang]
- Definition Classes
- AnyRef
- Annotations
- @throws(classOf[java.lang.Throwable]) @Deprecated
- Deprecated
(Since version 9)