Packages

class NameScope extends SQLConfHelper

The NameScope is used to control the resolution of names (table, column, alias identifiers). It's a part of the Resolver's state, and is used to manage the output of SQL query/DataFrame program operators.

The NameScope output is immutable. If it's necessary to update the output, NameScopeStack methods are used (overwriteCurrent or pushScope). The NameScope is always used through the NameScopeStack.

The resolution of identifiers is case-insensitive.

Name resolution priority is as follows:

  1. Resolution of local references:
    • column reference
    • parameterless function reference
    • struct field or map key reference 2. Resolution of lateral column aliases (if enabled). 3. In the context of Aggregate: resolution of names in grouping expressions list referencing aliases in aggregate expressions.

Following examples showcase the priority of name resolution:

SELECT 1 AS col1, col1 FROM VALUES (2)

Because column resolution has a higher priority than LCA resolution, the result will be [1, 2] and not [1, 1].

CREATE TABLE t AS SELECT col1 as current_date FROM VALUES (2);

SELECT
 1 AS current_timestamp,
 current_timestamp,
 current_date
FROM
 foo;

Result of the previous SELECT will be: [1, 2025-02-13T07:55:26.206+00:00, 2]. As can be seen, because of resolution precedence, current_date is resolved as a table column, but current_timestamp is resolved as a function without parenthesis instead of a lateral column reference.

Approximate tree of NameScope manipulations is shown in the following example:

CREATE TABLE IF NOT EXISTS t1 (col1 INT, col2 INT, col3 STRING);

SELECT
  col1, col2 as alias1
FROM
  (SELECT * FROM VALUES (1, 2))
  UNION
  (SELECT t2.col1, t2.col2 FROM (SELECT col1, col2 FROM t1) AS t2)
;

->

unionAttributes = pushScope {
  lhsOutput = pushScope {
    expandedStar = pushScope {
      scopes.overwriteCurrent(localRelation.output)
      scope.expandStar(star)
    }
    scopes.overwriteCurrent(expandedStar)
    scope.output
  }
  rhsOutput = pushScope {
    subqueryAttributes = pushScope {
      scopes.overwriteCurrent(t1.output)
      scopes.overwriteCurrent(prependQualifier(scope.output, "t2"))
      [scope.matchMultiPartName("t2", "col1"), scope.matchMultiPartName("t2", "col2")]
    }
    scopes.overwriteCurrent(subqueryAttributes)
    scope.output
  }
  scopes.overwriteCurrent(coerce(lhsOutput, rhsOutput))
  [scope.matchMultiPartName("col1"), alias(scope.matchMultiPartName("col2"), "alias1")]
}
scopes.overwriteCurrent(unionAttributes)
Linear Supertypes
SQLConfHelper, AnyRef, Any
Ordering
  1. Alphabetic
  2. By Inheritance
Inherited
  1. NameScope
  2. SQLConfHelper
  3. AnyRef
  4. Any
  1. Hide All
  2. Show All
Visibility
  1. Public
  2. Protected

Instance Constructors

  1. new NameScope(output: Seq[Attribute] = Seq.empty, hiddenOutput: Seq[Attribute] = Seq.empty, isSubqueryRoot: Boolean = false, availableAliases: HashSet[ExprId] = new HashSet[ExprId], aggregateListAliases: Seq[Alias] = Seq.empty, baseAggregate: Option[Aggregate] = None, planLogger: PlanLogger = new PlanLogger)

    output

    These are the attributes visible for lookups in the current scope. These may be:

    • Transformed outputs of lower scopes (e.g. type-coerced outputs of Union's children).
    • Output of a current operator that is being resolved (leaf nodes like Relations).
    hiddenOutput

    Attributes that are not directly visible in the scope, but available for lookup in case the resolved attribute is not found in output.

    isSubqueryRoot

    Indicates that the current scope is a root of a subquery. This is used by NameScopeStack.resolveMultipartName to detect the nearest outer scope.

    availableAliases

    User specified aliases that are present in this NameScope.

    aggregateListAliases

    List of aliases that are present in the Aggregate corresponding to this NameScope. If the Aggregate has lateral column references, this list contains both the aliases from Aggregate as well as all aliases from artificially inserted Project nodes.

    baseAggregate

    Aggregate node that is either a resolved Aggregate corresponding to this node or base Aggregate constructed when resolving lateral column references in Aggregate.

    planLogger

    PlanLogger used to log name resolution events.

Value Members

  1. final def !=(arg0: Any): Boolean
    Definition Classes
    AnyRef → Any
  2. final def ##: Int
    Definition Classes
    AnyRef → Any
  3. final def ==(arg0: Any): Boolean
    Definition Classes
    AnyRef → Any
  4. def addTopAggregateExpression(aliasedAggregateExpression: Alias): Unit

    Add a top level alias to the map so it can be used when resolving a grouping expression.

    Add a top level alias to the map so it can be used when resolving a grouping expression. We store Aliases with a same name in a list related to it as the order of the definition of the aliases in the aggregate expressions list is important (first one should be used for resolution in group by alias case). Example:

    SELECT col1 as a, col2 AS a FROM values('a', 'b') GROUP BY col2, a;

    a from the grouping expressions should be resolved as col1 (as it comes before col2 AS a in the aggregate expressions list). Plan looks like:

    Aggregate [col2#2, col1#1], [col1#1 AS a#3, col2#2 AS a#4] +- LocalRelation [col1#1, col2#2]

  5. val aggregateListAliases: Seq[Alias]
  6. final def asInstanceOf[T0]: T0
    Definition Classes
    Any
  7. val availableAliases: HashSet[ExprId]
  8. val baseAggregate: Option[Aggregate]
  9. def clone(): AnyRef
    Attributes
    protected[lang]
    Definition Classes
    AnyRef
    Annotations
    @throws(classOf[java.lang.CloneNotSupportedException]) @IntrinsicCandidate() @native()
  10. def conf: SQLConf

    The active config object within the current scope.

    The active config object within the current scope. See SQLConf.get for more information.

    Definition Classes
    SQLConfHelper
  11. final def eq(arg0: AnyRef): Boolean
    Definition Classes
    AnyRef
  12. def equals(arg0: AnyRef): Boolean
    Definition Classes
    AnyRef → Any
  13. def expandStar(unresolvedStar: Star): Seq[NamedExpression]

    Expand the Star.

    Expand the Star. The expected use case for this method is star expansion inside Project.

    Star without a target:

    -- Here the star will be expanded to [a, b, c].
    SELECT * FROM VALUES (1, 2, 3) AS t(a, b, c);

    Star with a multipart name target:

    USE CATALOG catalog1;
    USE DATABASE database1;
    
    CREATE TABLE IF NOT EXISTS table1 (col1 INT, col2 INT);
    
    -- Here the star will be expanded to [col1, col2].
    SELECT catalog1.database1.table1.* FROM catalog1.database1.table1;

    Star with a struct target:

    -- Here the star will be expanded to [field1, field2].
    SELECT d.* FROM VALUES (named_struct('field1', 1, 'field2', 2)) AS t(d);

    Star as an argument to a function:

    -- Here the star will be expanded to [col1, col2, col3] and those would be passed as
    -- arguments to `concat_ws`.
    SELECT concat_ws('', *) AS result FROM VALUES (1, 2, 3);

    Also, see Star.expandStar for more details.

  14. def findAttributesByName(name: String): Seq[Attribute]

    Find attributes in this NameScope that match a provided one-part name.

    Find attributes in this NameScope that match a provided one-part name.

    This method is simpler and more lightweight than resolveMultipartName, because here we just return all the attributes matched by the one-part name. This is only suitable for situations where name _resolution_ is not required (e.g. accessing struct fields from the lower operator's output).

    For example, this method is used to look up attributes to match a specific View schema. See ExpressionResolver.resolveGetViewColumnByNameAndOrdinal for more info on view column lookup.

    We are relying on a simple IdentifierMap to perform that work, since we just need to match one-part name from the lower operator's output here.

  15. def getAttributeById(expressionId: ExprId): Option[Attribute]

    Returns attribute with expressionId if output contains it.

    Returns attribute with expressionId if output contains it. This is used to preserve nullability for resolved AttributeReference.

  16. final def getClass(): Class[_ <: AnyRef]
    Definition Classes
    AnyRef → Any
    Annotations
    @IntrinsicCandidate() @native()
  17. def getHiddenAttributeById(expressionId: ExprId): Option[Attribute]

    Returns attribute with expressionId if hiddenOutput contains it.

  18. def getOrdinalReplacementExpressions: Option[OrdinalReplacementExpressions]

    Get expressions that are candidates to be resolved by ordinal in current NameScope.

  19. def getOutputIds: Set[ExprId]

    Return all the explicitly outputted expression IDs.

    Return all the explicitly outputted expression IDs. Hidden or metadata output are not included.

  20. def getTopAggregateExpressionAliases: Seq[Alias]

    Returns all the top-level aliases from Aggregate corresponding to this NameScope collected while resolving aggregate expressions.

    Returns all the top-level aliases from Aggregate corresponding to this NameScope collected while resolving aggregate expressions. Used to resolve group by alias.

  21. def hashCode(): Int
    Definition Classes
    AnyRef → Any
    Annotations
    @IntrinsicCandidate() @native()
  22. val hiddenOutput: Seq[Attribute]
  23. final def isInstanceOf[T0]: Boolean
    Definition Classes
    Any
  24. val isSubqueryRoot: Boolean
  25. lazy val lcaRegistry: LateralColumnAliasRegistry
  26. final def ne(arg0: AnyRef): Boolean
    Definition Classes
    AnyRef
  27. final def notify(): Unit
    Definition Classes
    AnyRef
    Annotations
    @IntrinsicCandidate() @native()
  28. final def notifyAll(): Unit
    Definition Classes
    AnyRef
    Annotations
    @IntrinsicCandidate() @native()
  29. val output: Seq[Attribute]
  30. def overwrite(output: Option[Seq[Attribute]] = None, hiddenOutput: Option[Seq[Attribute]] = None, availableAliases: Option[HashSet[ExprId]] = None, aggregateListAliases: Seq[Alias] = Seq.empty, baseAggregate: Option[Aggregate] = None): NameScope

    Returns new NameScope which preserves all the immutable NameScope properties but overwrites output, hiddenOutput, availableAliases, aggregateListAliases and baseAggregate if provided.

    Returns new NameScope which preserves all the immutable NameScope properties but overwrites output, hiddenOutput, availableAliases, aggregateListAliases and baseAggregate if provided. Mutable state like lcaRegistry is not preserved.

  31. def resolveMissingAttributesByHiddenOutput(referencedAttributes: HashMap[ExprId, Attribute]): Seq[Attribute]

    Given referenced attributes, returns all attributes that are referenced and missing from current output, but can be found in hidden output.

  32. def resolveMultipartName(multipartName: Seq[String], canLaterallyReferenceColumn: Boolean = false, canReferenceAggregateExpressionAliases: Boolean = false, canResolveNameByHiddenOutput: Boolean = false, shouldPreferHiddenOutput: Boolean = false, canReferenceAggregatedAccessOnlyAttributes: Boolean = false): NameTarget

    Resolve multipart name into a NameTarget.

    Resolve multipart name into a NameTarget. NameTarget's candidates may contain simple AttributeReferences if it's a column or alias, or ExtractValue expressions if it's a struct field, map value or array value. The aliasName will optionally be set to the proposed alias name for the value extracted from a struct, map or array.

    Example that demonstrates those major use-cases:

    CREATE TABLE IF NOT EXISTS t (
      col1 INT,
      col2 STRUCT<field: INT>,
      col3 STRUCT<struct: STRUCT<field: INT>>,
      col4 MAP<STRING: INT>,
      col5 STRING
    );
    
    -- For the SELECT below the top Project list will be resolved using this method like this:
    -- AttributeReference(col1),
    -- AttributeReference(a),
    -- GetStructField(col2, field),
    -- GetStructField(GetStructField(col3, struct), field),
    -- GetMapValue(col4, key)
    SELECT
      col1, a, col2.field, col3.struct.field, col4.key
    FROM
      (SELECT *, col5 AS a FROM t);

    Since there can be several expressions that matched the same multipart name, this method may return a NameTarget with the following candidates: - 0 values: No matched expressions - 1 value: Unique expression matched - 1+ values: Ambiguity, several expressions matched

    Some examples of ambiguity:

    CREATE TABLE IF NOT EXISTS t1 (c1 INT, c2 INT);
    CREATE TABLE IF NOT EXISTS t2 (c2 INT, c3 INT);
    
    -- Identically named columns from different tables.
    -- This will fail with AMBIGUOUS_REFERENCE error.
    SELECT c2 FROM t1, t2;
    CREATE TABLE IF NOT EXISTS foo (c1 INT);
    CREATE TABLE IF NOT EXISTS bar (foo STRUCT<c1: INT>);
    
    -- Ambiguity between a column in a table and a field in a struct.
    -- This will succeed, and column will win over the struct field.
    SELECT foo.c1 FROM foo, bar;

    The candidates are deduplicated by expression ID (not by attribute name!):

    CREATE TABLE IF NOT EXISTS t1 (col1 STRING);
    
    -- No ambiguity here, since we are selecting the same column (same expression ID).
    SELECT col1 FROM (SELECT col1, col1 FROM t);

    The case of the multipartName takes precedence over the original name case, so the candidates will have names that are case-identical to the multipartName:

    CREATE TABLE IF NOT EXISTS t1 (col1 STRING);
    
    -- The output schema of this query is [COL1], despite the fact that the column is in
    -- lower-case.
    SELECT COL1 FROM t;

    Name resolution can be done using the hidden output for certain operators (e.g Sort, Filter). This is indicated by canResolveNameByHiddenOutput which is passed from ExpressionResolver.resolveAttribute based on the parent operator. Example:

    -- Project's output = [`col1`]; Project's hidden output = [`col1`, `col2`]
    SELECT col1 FROM VALUES(1, 2) ORDER BY col2;

    When resolving Sort, SortOrder expressions should first be attempted to be resolved by table columns in project list and only then by aliases from the project list. For example, in the following query:

    SELECT 1 AS col1, col1 FROM VALUES(1) ORDER BY col1

    Even though there is ambiguity with the name col1, the SortOrder expression should be resolved as a table column from the project list and not throw AMBIGUOUS_REFERENCE.

    The names in Aggregate.groupingExpressions can reference Aggregate.aggregateExpressions aliases. canReferenceAggregateExpressionAliases will be true when we are resolving the grouping expressions. Example:

    {{ SELECT col1 + col2 AS a FROM VALUES (1, 2) GROUP BY a; }}}

    In case we are resolving names in expression trees from HAVING or ORDER BY on top of Aggregate, we are able to resolve hidden attributes only if those are present in grouping expressions, or if the reference itself is under an AggregateExpression. In the latter case canReferenceAggregatedAccessOnlyAttributes will be true, and all the attributes from hiddenOutput will be used instead of the former case where we don't use attributes marked as aggregatedAccessOnly. Consider the following example:

    -- This succeeds, because `col2` is in the grouping expressions.
    SELECT COUNT(col1) FROM t1 GROUP BY col1, col2 ORDER BY col2;
    
    -- This fails, because `col2` is not in the grouping expressions.
    SELECT COUNT(col1) FROM t1 GROUP BY col1 ORDER BY col2;
    
    -- This succeeds, despite the fact that `col2` is not in the grouping expressions.
    -- Such references are allowed under an aggregate expression (MAX).
    SELECT COUNT(col1) FROM t1 GROUP BY col1 ORDER BY MAX(col2);

    Spark is being smart about name resolution and prioritizes candidates from output levels that can actually be resolved, even though that output level might not be the first choice. For example, ORDER BY clause prefers attributes from SELECT list (namely, aliases) over table columns from below. However, if attributes on the SELECT level have name ambiguity or other issues, Spark will try to resolve the name using the table columns from below. Examples:

    CREATE TABLE t1 (col1 INT);
    CREATE TABLE t2 (col1 STRUCT<field: INT>);
    
    -- Main output is ambiguous, so col1 from t1 is used for sorting.
    SELECT 1 AS col1, 2 AS col1 FROM t1 ORDER BY col1;
    
    -- col1 from main output does not have `field`, so struct field of col1 from t2 is used for
    -- sorting.
    SELECT 1 AS col1 FROM t2 ORDER BY col1.field;

    This is achieved using candidate prioritization mechanism in pickSuitableCandidates.

    We are relying on the AttributeSeq to perform name resolution, since it requires complex resolution logic involving nested field extraction and multipart name matching. See AttributeSeq.resolve for more details.

  33. def setOrdinalReplacementExpressions(ordinalReplacementExpressions: OrdinalReplacementExpressions): Unit

    Set expressions that are candidates to be resolved by ordinal in current NameScope.

  34. final def synchronized[T0](arg0: => T0): T0
    Definition Classes
    AnyRef
  35. def toString(): String
    Definition Classes
    AnyRef → Any
  36. final def wait(arg0: Long, arg1: Int): Unit
    Definition Classes
    AnyRef
    Annotations
    @throws(classOf[java.lang.InterruptedException])
  37. final def wait(arg0: Long): Unit
    Definition Classes
    AnyRef
    Annotations
    @throws(classOf[java.lang.InterruptedException]) @native()
  38. final def wait(): Unit
    Definition Classes
    AnyRef
    Annotations
    @throws(classOf[java.lang.InterruptedException])
  39. def withSQLConf[T](pairs: (String, String)*)(f: => T): T

    Sets all SQL configurations specified in pairs, calls f, and then restores all SQL configurations.

    Sets all SQL configurations specified in pairs, calls f, and then restores all SQL configurations.

    Attributes
    protected
    Definition Classes
    SQLConfHelper

Deprecated Value Members

  1. def finalize(): Unit
    Attributes
    protected[lang]
    Definition Classes
    AnyRef
    Annotations
    @throws(classOf[java.lang.Throwable]) @Deprecated
    Deprecated

    (Since version 9)

Inherited from SQLConfHelper

Inherited from AnyRef

Inherited from Any

Ungrouped