Skip to content

Preserve nested column resolution in AvroSchemaUtils.toAvroSchema #375

Description

@coderabbitai

Summary

AvroSchemaUtils.toAvroSchema accepts a sequence of column names. Its current implementation builds a StructType from dataFrame.schema(columnName). This lookup only resolves top-level schema fields.

Investigate and define the intended behavior for nested column paths such as Seq("parent.child"). If nested paths must be supported, preserve Catalyst column resolution when building the schema.

Rationale

This behavior is outside the Spark 4 upgrade scope. A change can alter behavior for existing callers, so it requires separate investigation and regression coverage.

Affected area

  • src/main/scala/za/co/absa/abris/avro/parsing/utils/AvroSchemaUtils.scala
  • Spark schema conversion tests for AvroSchemaUtils

Required work

  1. Identify existing callers that pass nested column paths to the columnNames: Seq[String] overload.
  2. Define the compatibility contract for nested paths.
  3. If nested paths are supported, construct the selected schema through analyzed DataFrame column expressions rather than direct top-level StructType field lookup.
  4. Add regression tests for a nested path such as Seq("parent.child").
  5. Document any intentional behavior change or compatibility limitation.

Acceptance criteria

  • The nested-column behavior is explicitly defined.
  • Tests cover the defined behavior on the supported Spark 4 version.
  • The implementation does not unintentionally change top-level column behavior.

Backlinks

Requested by: @kevinwallimann

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Labels

No labels
No labels

Type

No type

Projects

No projects

    Milestone

    No milestone

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions