Packages

t

org.apache.spark.sql.catalyst.analysis.resolver

ResolvesNameByHiddenOutput

trait ResolvesNameByHiddenOutput extends SQLConfHelper

ResolvesNameByHiddenOutput is used by resolvers for operators that are able to resolve attributes in its expression tree from hidden output or that can reference expressions not present in child's output. Update child operator's output list and place a Project node on top of original operator node with the original output of an operator's child.

For example, in a following query:

SELECT
 t1.key
FROM
 t1 FULL OUTER JOIN t2 USING (key)
WHERE
 t1.key NOT LIKE 'bb.%';

Plan without adding missing attributes would be:

+- Project [key#1]
   +- Filter NOT key#1 LIKE bb.%
      +- Project [coalesce(key#1, key#2) AS key#3, __key#1__, __key#2__]
         +- Join FullOuter, (key#1 = key#2)
            :- SubqueryAlias t1
            :  +- Relation t1[key#1]
            +- SubqueryAlias t2
               +- Relation t2[key#2]

NOTE: #key1 and key#2 at the end of inner Project are metadata columns from the full outer join. Even though they are present in the Project in single-pass, fixed-point adds these columns after resolving missing input, so duplication of some columns is possible. In order to stay fully compatible between single-pass and fixed-point, we add both missing attributes and these metadata columns. We mimic fixed-point behavior by putting metadata columns in NameScope.hiddenOutput instead of NameScope.output.

In the plan above, Filter requires key#1 in its condition, but key#1 is not available in the below Project's output, even though key#1 is available in Join's hidden output. Because of that, we need to place key#1 in the project list, after original project list expressions, but before metadata columns (to remain compatible with fixed-point). In order to preserve initial output of Filter, we place a Project node on top of this Filter, whose project list is the original output of the Project below Filter (in this case - key#3 and metadata columns key#1 and key#2).

Therefore, the plan becomes:

+- Project [key#1]
   +- Project [key#3, key#1, key#2]
      +- Filter NOT key#1 LIKE bb.%
         +- Project [coalesce(key#1, key#2) AS key#3, key#1, key#1, key#2]
            +- Join FullOuter, (key#1 = key#2)
               :- SubqueryAlias t1
               :  +- Relation t1[key#1]
               +- SubqueryAlias t2
                  +- Relation t2[key#2]

Query below exhibits similar behavior when Sort operator resolves an attribute using hidden output:

SELECT col1 FROM VALUES (1, 2) ORDER BY col2;

Unresolved plan would be:

Sort [col2 ASC NULLS FIRST], true
  +- Project [col1]
    +- LocalRelation [col1, col2]

As it can be seen, attribute col2 used in Sort can't be resolved using the Project output (which is [col1]), so it has to be resolved using the hidden output (which is propagated from LocalRelation and is [col1, col2]). As it's been shown in the previous example, col2 has to be added to Project list and a Project with original output of the Project below Sort is added as a top node. Because of that, analyzed plan is:

Project [col1]
  +- Sort [col2 ASC NULLS FIRST], true
    +- Project [col1, col2]
      +- LocalRelation [col1, col2]

Another example is when Sort order expression is an AggregateExpression which is not present in the Aggregate.aggregateExpressions:

SELECT col1 FROM VALUES (1) GROUP BY col1 ORDER BY sum(col1);

In this example sum(col1) should be added to child's output and a Project node should be added on top of the Sort node to preserve the original output of the Aggregate node:

Project [col1] +- Sort [sum(col1)#... ASC NULLS FIRST], true +- Aggregate [col1], [col1, sum(col1) AS sum(col1)#...] +- LocalRelation [col1]

In case of Dataframe programs we can have multiple Sort operators nested inside each other. For example:

df.select("col1").orderBy("col2").orderBy("col1").orderBy("col2")

Unresolved plan would be:

Sort [col2 ASC NULLS FIRST], true
  +- Sort [col1 ASC NULLS FIRST], true
    +- Project [col1]
      +- Sort [col2 ASC NULLS FIRST], true
        +- Project [col1, col2]
          +- Project [col1, col2, col3]
            +- LocalRelation [col1, col2, col3]

As it can be seen, col2 (Sort order expression) needs to be resolved using the hidden output. Because of that it must be added to all the Projects and Aggregates below the Sort operator. Resolved plan would be:

Project [col1]
  +- Sort [col2 ASC NULLS FIRST], true
    +- Sort [col1 ASC NULLS FIRST], true
      +- Project [col1, col2]
        +- Sort [col2 ASC NULLS FIRST], true
          +- Project [col1, col2]
            +- Project [col1, col2, col3]
              +- LocalRelation [col1, col2, col3]

In the plan you can see that col2 is added to the lower Project.projectList.

Linear Supertypes
SQLConfHelper, AnyRef, Any
Ordering
  1. Alphabetic
  2. By Inheritance
Inherited
  1. ResolvesNameByHiddenOutput
  2. SQLConfHelper
  3. AnyRef
  4. Any
  1. Hide All
  2. Show All
Visibility
  1. Public
  2. Protected

Value Members

  1. final def !=(arg0: Any): Boolean
    Definition Classes
    AnyRef → Any
  2. final def ##: Int
    Definition Classes
    AnyRef → Any
  3. final def ==(arg0: Any): Boolean
    Definition Classes
    AnyRef → Any
  4. final def asInstanceOf[T0]: T0
    Definition Classes
    Any
  5. def clone(): AnyRef
    Attributes
    protected[lang]
    Definition Classes
    AnyRef
    Annotations
    @throws(classOf[java.lang.CloneNotSupportedException]) @IntrinsicCandidate() @native()
  6. def conf: SQLConf

    The active config object within the current scope.

    The active config object within the current scope. See SQLConf.get for more information.

    Definition Classes
    SQLConfHelper
  7. def deduplicateMissingExpressions(missingExpressions: Seq[NamedExpression]): Seq[NamedExpression]

    Deduplicates missing expressions by ExprId.

  8. final def eq(arg0: AnyRef): Boolean
    Definition Classes
    AnyRef
  9. def equals(arg0: AnyRef): Boolean
    Definition Classes
    AnyRef → Any
  10. final def getClass(): Class[_ <: AnyRef]
    Definition Classes
    AnyRef → Any
    Annotations
    @IntrinsicCandidate() @native()
  11. def hashCode(): Int
    Definition Classes
    AnyRef → Any
    Annotations
    @IntrinsicCandidate() @native()
  12. def insertMissingExpressions(operator: LogicalPlan, missingExpressions: Seq[NamedExpression]): LogicalPlan

    Insert the missing expressions in the output list of the operator.

    Insert the missing expressions in the output list of the operator. Recursively call expandOperatorsOutputList to expand the output lists of Projects and Aggregates below the current one. In order to stay compatible with fixed-point, missing expressions are inserted after the original output list, but before any qualified access only columns that have been added as part of resolution from hidden output.

    Only AttributeReferences are propagated recursively in expandOperatorsOutputList. Aliases are meant to be inserted in the topmost operator. For example, SortResolver may push down aliased aggregate and grouping expression trees to the immediate Aggregate below, but it does not make sense to push them down further:

    -- The MAX(v2.col1) aggregate has to be pushed down to [[Aggregate]], but not to [[Project]]
    -- below it.
    SELECT COUNT(col1) FROM v1 NATURAL JOIN v2 GROUP BY col1 ORDER BY MAX(v2.col1);

    ->

    Project [count(col1)#23L]
    +- Sort [max(col1)#25 ASC NULLS FIRST], true
      +- Aggregate [col1#21], [count(col1#21) AS count(col1)#23L, max(col1#20) AS max(col1)#25]
        +- Project [col1#21, col1#20]
          ...
  13. final def isInstanceOf[T0]: Boolean
    Definition Classes
    Any
  14. final def ne(arg0: AnyRef): Boolean
    Definition Classes
    AnyRef
  15. final def notify(): Unit
    Definition Classes
    AnyRef
    Annotations
    @IntrinsicCandidate() @native()
  16. final def notifyAll(): Unit
    Definition Classes
    AnyRef
    Annotations
    @IntrinsicCandidate() @native()
  17. def retainOriginalOutput(operator: LogicalPlan, missingExpressions: Seq[NamedExpression], scopes: NameScopeStack): LogicalPlan

    If missingExpressions is not empty, output of an operator has been changed by insertMissingExpressions.

    If missingExpressions is not empty, output of an operator has been changed by insertMissingExpressions. Therefore, we need to restore the original output, by placing a Project on top of an original node, with original's node output. Additionally, we append all qualified access only columns from hidden output that were inserted as missing attributes, because they may be needed in upper operators (if not, they will be pruned away in PruneMetadataColumns). Other hidden attributes are thrown away, because we cannot reference them from the new Project (they are not outputted from below).

    If SQLConf.SINGLE_PASS_RESOLVER_PREVENT_USING_ALIASES_FROM_NON_DIRECT_CHILDREN is set to true, we need to overwrite the current scope and clear aggregateListAliases and baseAggregate. This is needed in order to prevent later replacement of Sort/Having expressions using semantically equal aliased expressions from non-direct children. For example, in the following query:

    SELECT col1 AS a FROM VALUES(1,2) GROUP BY col1, col2 HAVING col2 > 1 ORDER BY col1;

    With flag set to false, analyzed plan will be:

    Sort [a#3 ASC NULLS FIRST], true +- Project [a#3] +- Filter (col2#2 > 1) +- Aggregate [col1#1, col2#2], [col1#1 AS a#3, col2#2, col1#1] +- LocalRelation [col1#1, col2#2]

    Instead of using missing attribute col1#1 we can use its alias a#3 in the Sort and avoid adding an extra projection. This is because all of Sort, Project, Filter and Aggregate belong to the same NameScope since Project was artificially inserted.

    However, fixed-point can't handle this case properly and produces the following plan:

    Project [a#3] +- Sort [col1#1 ASC NULLS FIRST], true +- Project [a#3, col1#1] +- Filter (col2#2 > 1) +- Aggregate [col1#1, col2#2], [col1#1 AS a#3, col2#2, col1#1] +- LocalRelation [col1#1, col2#2]

    Therefore, we need to match this behavior of fixed-point in single-pass in order to avoid logical plan mismatches.

  18. final def synchronized[T0](arg0: => T0): T0
    Definition Classes
    AnyRef
  19. def toString(): String
    Definition Classes
    AnyRef → Any
  20. final def wait(arg0: Long, arg1: Int): Unit
    Definition Classes
    AnyRef
    Annotations
    @throws(classOf[java.lang.InterruptedException])
  21. final def wait(arg0: Long): Unit
    Definition Classes
    AnyRef
    Annotations
    @throws(classOf[java.lang.InterruptedException]) @native()
  22. final def wait(): Unit
    Definition Classes
    AnyRef
    Annotations
    @throws(classOf[java.lang.InterruptedException])
  23. def withSQLConf[T](pairs: (String, String)*)(f: => T): T

    Sets all SQL configurations specified in pairs, calls f, and then restores all SQL configurations.

    Sets all SQL configurations specified in pairs, calls f, and then restores all SQL configurations.

    Attributes
    protected
    Definition Classes
    SQLConfHelper

Deprecated Value Members

  1. def finalize(): Unit
    Attributes
    protected[lang]
    Definition Classes
    AnyRef
    Annotations
    @throws(classOf[java.lang.Throwable]) @Deprecated
    Deprecated

    (Since version 9)

Inherited from SQLConfHelper

Inherited from AnyRef

Inherited from Any

Ungrouped