Skip to content

[736] Make the bundled xtable-utilities jar runnable again (Hudi 0.x) + CI on branch-0.4 - #841

Merged
vinishjail97 merged 3 commits into
apache:branch-0.4from
vinishjail97:736-jol-utilities-hudi0x
Jul 20, 2026
Merged

[736] Make the bundled xtable-utilities jar runnable again (Hudi 0.x) + CI on branch-0.4#841
vinishjail97 merged 3 commits into
apache:branch-0.4from
vinishjail97:736-jol-utilities-hudi0x

Conversation

@vinishjail97

@vinishjail97 vinishjail97 commented Jul 9, 2026

Copy link
Copy Markdown
Contributor

Backports #840 to the Hudi 0.x line (branch-0.4) for the 0.4.0 release branch.

Closes #736

vinishjail97 and others added 2 commits July 9, 2026 11:38
Cherry-pick of apache#840 (4daec27) adapted for the Hudi 0.14 line on
main-hudi-0x. The apache#736 root cause applies here too: the apache#822 shade
allowlist dropped runtime deps from the bundled utilities jar, so
java -jar RunSync failed with NoClassDefFoundError.

Applicable subset for Hudi 0.x:
- jol-core: the parent pins it to test scope, but Hudi's
  ObjectSizeCalculator loads org.openjdk.jol at runtime; override to
  runtime and add to the shade allowlist (the actual apache#736 fix).
- slf4j-api: add to the allowlist (org/slf4j/LoggerFactory was missing).

Dropped from the original apache#840 (Hudi 1.x only): hudi-hadoop-common and
hudi-io do not exist in Hudi 0.14 (that split is 1.x); hudi-common still
provides org.apache.hudi.common.fs.FSUtils. There is also no redundant
test-scoped hudi-java-client to remove on this branch.

Verified: the 0.x bundled jar now contains org/slf4j/LoggerFactory,
org/openjdk/jol/info/GraphLayout, org/apache/hudi/common/fs/FSUtils and
org/apache/hudi/client/common/HoodieJavaEngineContext.

Cherry picked from commit 4daec27.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Add main-hudi-0x to the push and pull_request branch triggers so the
existing CI (build + test, license check) runs for the long-lived
Hudi 0.x branch and PRs targeting it, matching main. package-deploy
(release-triggered) and the site workflows are unchanged.

The concurrency guard already treats any ref containing "main" as
non-cancelable, so main-hudi-0x is covered without further changes.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
@vinishjail97
vinishjail97 marked this pull request as ready for review July 9, 2026 18:51
vinishjail97 added a commit to vinishjail97/onetable that referenced this pull request Jul 10, 2026
Add main-hudi-0x to the push and pull_request branch filters of both
workflows (same change as apache#841) so CI runs on PRs targeting the Hudi 0.x
release line, including this one. A pull_request workflow's branch filter
takes effect from the PR's own branch, so this makes CI fire on apache#843.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
The release branch was renamed main-hudi-0x -> branch-0.4 to use a
version-keyed name instead of a dependency-keyed one. Update the push /
pull_request triggers so Maven CI Build and License Check run on the
renamed branch.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
@vinishjail97
vinishjail97 force-pushed the 736-jol-utilities-hudi0x branch from 46823b4 to 043c82b Compare July 13, 2026 15:54
@vinishjail97 vinishjail97 changed the title [736] Make the bundled xtable-utilities jar runnable again (Hudi 0.x) + run CI on main-hudi-0x [736] Make the bundled xtable-utilities jar runnable again (Hudi 0.x) + run CI on branch-0.4 Jul 13, 2026
@vinishjail97 vinishjail97 changed the title [736] Make the bundled xtable-utilities jar runnable again (Hudi 0.x) + run CI on branch-0.4 [736] Make the bundled xtable-utilities jar runnable again (Hudi 0.x) + CI on main-hudi-0x Jul 16, 2026
@vinishjail97 vinishjail97 changed the title [736] Make the bundled xtable-utilities jar runnable again (Hudi 0.x) + CI on main-hudi-0x [736] Make the bundled xtable-utilities jar runnable again (Hudi 0.x) + CI on branch-0.4 Jul 16, 2026

@vinothchandar vinothchandar left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Makes sense. lgtm

@vinishjail97
vinishjail97 merged commit 7e273e6 into apache:branch-0.4 Jul 20, 2026
2 checks passed
@vinishjail97
vinishjail97 deleted the 736-jol-utilities-hudi0x branch July 20, 2026 17:22
vinishjail97 added a commit that referenced this pull request Jul 21, 2026
…k 3.4 + 3.5) (#843)

* [836] xtable-spark-runtime: thin drop-in Spark bundle (Hudi 0.x)

Ports the xtable-spark-runtime packaging slice onto the Hudi 0.x line
(main-hudi-0x, Hudi 0.14 / Spark 3.4). New thin, relocated bundle that
runs an incremental XTable sync inside a Spark job; engines are provided
by the cluster, never bundled.

Module:
- xtable-spark-runtime_${scala.binary.version}: xtable-core compile;
  Spark/Hadoop and the engines (Hudi/Iceberg/Delta) provided. Curated
  shade allowlist (xtable modules + guava/protobuf/commons-cli relocated);
  avro/parquet/jackson NOT relocated (exchanged with the engines). Thin
  ~3.7 MB bundle.
- XTableSparkSync: standalone spark-submit entry point (Apache Commons CLI).
- XTableSyncService / TableSyncSpec: build an INCREMENTAL ConversionConfig
  and run ConversionController.sync; target metadata path = source data
  path (required by Hudi; Iceberg data lives under <basePath>/data).

Hudi 0.14 specifics (vs the Hudi 1.x variant on main):
- No hudi-hadoop-common (that split is 1.x); hudi-common provides FSUtils.
- Engine classpath uses the hudi-spark bundle (its regenerated Avro model
  classes link on Avro 1.12, which Iceberg 1.9.2 requires) plus
  hudi-java-client for the Hudi target's Java write client, with the raw
  hudi-common excluded so the bundle's clean DecimalWrapper wins.

ConversionTargetFactory: make ServiceLoader discovery resilient so a
subset of engines works when others are absent (warn + skip on
LinkageError/ServiceConfigurationError; name-based Delta-Kernel check).

ITXTableSparkRuntimeBundle: spark-submits the shaded jar for one case per
direction across Hudi, Iceberg and Delta (source and target), engines on
a flat classpath, asserting data-equivalence over comparable scalar
columns. Requires a Spark 3.4 SPARK_HOME; skipped otherwise.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* [836] xtable-spark-runtime: address review feedback (packaging, licensing, RFC)

Publish the shaded jar as the MAIN artifact with a dependency-reduced POM
(shadedArtifactAttached=false, createDependencyReducedPom=true), matching
iceberg-spark-runtime / hudi-spark-bundle. Resolving the coordinate via a
Maven dependency or --packages now yields the relocated bundle and pulls no
un-relocated transitive deps; --jars is equivalent. The thin main jar +
separate -bundle classifier (which re-introduced the cluster guava clash)
is gone. IT findBundleJar() now picks the shaded main jar.

Pass release/scripts/validate_shaded_license_coverage.sh (the existing
allowlist gate): the shade <includes> must equal the runtime dependency
tree, so jackson / scala-library / log4j-1.2-api - which every Spark
runtime supplies - are declared provided (dropped from the tree, kept off
the shaded jar) and excluded from the IT's flat engine classpath so Spark's
own copies win. avro/parquet stay on the flat classpath (the engine's newer
avro must win over Spark 3.4's).

Bundled-dependency licensing: add META-INF/LICENSE-bundled and
NOTICE-bundled (wired via IncludeResourceTransformer) attributing the only
bundled third-party - guava's closure (Apache-2.0) and commons-cli
(Apache-2.0), plus checker-qual (MIT). Remove the dead protobuf-java
allowlist entry and relocation: protobuf resolves as provided and was never
bundled.

RFC-3: scope v1 to the CLI (XTableSparkSync); mark the config-only
XTableSyncListener on-ramp (and XTableSparkConfig / PlanTargetResolver) as a
deferred follow-up. Fix the activation example to the shipped CLI and
document the required engine Avro version + flat-classpath placement.

XTableSyncService: normalize sourceFormat with Locale.ROOT.

Root pom: exclude the shade-generated dependency-reduced-pom.xml from
spotless (no license header; CI runs clean install).

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* [836] xtable-spark-runtime: trim duplicated pom comments

Consolidate repeated rationale in the bundle pom (no functional change):
- Tell the hudi-common DecimalWrapper / Avro-1.8-1.9 story once (on the
  hudi-spark bundle dep); the hudi-java-client exclusion just points to it.
- State "engine Avro must win on a flat classpath" once (engine-classpath
  plugin); the avro dep and excludeGroupIds comments reference it.
- Explain jackson/scala-library are Spark-supplied once.
- Fix a stale "Spark 3.5" reference to Spark 3.4 (this is the Hudi 0.x line).

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* [836] ci: trigger Maven CI + License Check on main-hudi-0x

Add main-hudi-0x to the push and pull_request branch filters of both
workflows (same change as #841) so CI runs on PRs targeting the Hudi 0.x
release line, including this one. A pull_request workflow's branch filter
takes effect from the PR's own branch, so this makes CI fire on #843.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* Address review feedback: guard ServiceLoader.hasNext() and drain spark-submit stdout off-thread

- ConversionTargetFactory: hasNext() resolves provider classes lazily and can
  throw ServiceConfigurationError too, so move it inside the existing
  ServiceConfigurationError|LinkageError guard alongside next().
- ITXTableSparkRuntimeBundle: drain spark-submit stdout on a background thread
  so a hung process is caught by waitFor(10min) instead of blocking forever on
  readOutput() reading until EOF.

* [836] ci: point Maven CI + License Check triggers at branch-0.4

main-hudi-0x was deleted and replaced by branch-0.4 (the 0.4.0 release line),
so update the push/pull_request branch filters accordingly.

* [836] ci: validate xtable-spark-runtime bundle against a real Spark distro

ITXTableSparkRuntimeBundle spark-submits the shaded bundle jar and
self-skips unless SPARK_HOME is set, so the main CI never runs it. Add a
path-filtered workflow that installs a matching Spark 3.4 distribution
from the Apache archive, sets SPARK_HOME, and runs the failsafe IT.

* [836] ci: run spark-runtime validation on every branch PR (drop path filter)

The workflow is intended to be a required status check on branch-0.4. A
required check whose workflow is skipped by a path filter never reports,
leaving PRs blocked on a pending check, so run it unconditionally on the
covered branches. The Spark distro is cached, so the added cost is the
reactor build plus the ~20s IT.

* [836] xtable-spark-runtime: address review feedback (CLI validation, unix flags, docs)

- XTableSparkSync: validate --sourceformat/--targets up front and fail fast
  (before SparkSession creation) on empty or unsupported values. Split the
  allowed sets: sources are Hudi/Iceberg/Delta/Paimon/Parquet, targets are
  Hudi/Iceberg/Delta (Paimon/Parquet are read-only, no ConversionTarget).
- XTableSyncService.sourceProviderFor: wire Paimon and Parquet source providers
  (previously threw UnsupportedOperationException); engines remain cluster-provided.
- Rename CLI long-opts to unix-style lowercase (--basepath, --sourceformat,
  --tablename, --datapath, --partitionspec); update javadoc, IT and RFC example.
- basePathToName -> basePathToTableName: handle "/", trailing slashes and null
  by throwing with a "pass --tablename" hint instead of an empty table name.
- Add XTableSparkSyncTest covering table-name derivation and source/target
  format validation.
- ConversionTargetFactory: log the skipped provider's error class/message so
  operators can distinguish an intentionally-absent engine from a linkage error.
- spark-runtime-validation.yml: add a workflow_dispatch spark_version input and
  document why the Spark 3.4 line is pinned (Delta 2.4.0 is Spark-3.4-only).
- RFC-3: add a supported-formats/engine-versions section, "from application
  code" (Scala/Java + PySpark) activation examples, and list @vinothchandar as
  an approver.

* [836] xtable-spark-runtime: support Spark 3.5 via Delta Kernel

Delta source/target now auto-switch from delta-core to the Spark-free
Delta Kernel implementation on Spark 3.5+ (where the bundled delta-core
does not run), controlled by a single --usedeltakernel toggle that is
auto-enabled by Spark version. Hudi/Iceberg sync was already Spark-free.

Also runs the spark-runtime bundle IT on both Spark 3.4.3 and 3.5.9 via
a CI matrix.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* [836] Give Maven CI Build and License Check unique job names

Both workflows used a job named "build", so both reported the same
status-check context and could not be required distinctly. Set unique
job names so branch protection (see #848) can
require each on branch-0.4.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* [836] xtable-spark-runtime: add --datasetconfig for multi-table sync

XTableSparkSync accepts a --datasetconfig YAML (sourceFormat, targetFormats,
datasets[]) to sync multiple tables in one run, mutually exclusive with the
single-table --basepath/--sourceformat/--targets flags. The config is read
through the Spark Hadoop config, so it may live on a local or cloud
(s3/gcs/abfs) path, and reuses the same schema as the RunSync utility.

Parsed with SnakeYAML's SafeConstructor (plain maps/lists only, no arbitrary
type instantiation). SnakeYAML is bundled and relocated to
org.apache.xtable.shaded (like guava/commons-cli) because Spark ships its own
version (1.33 on 3.4, 2.0 on 3.5) that would otherwise clash; its multi-release
classes are filtered out so no un-relocated org.yaml classes remain.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants