Repository navigation
Conversation
…n spark built for hadoop 2.3.0 , 2.4.0
|
Can one of the admins verify this patch? |
Contributor
|
Jenkins, test this please. |
|
Merged build triggered. |
|
Merged build started. |
|
Merged build finished. All automated tests passed. |
|
All automated tests passed. |
Contributor
|
Thanks - I tested this locally. |
asfgit
pushed a commit
that referenced
this pull request
May 5, 2014
…n s... ...park built for hadoop 2.3.0 , 2.4.0 Author: witgo <[email protected]> Closes #628 from witgo/SPARK-1693_new and squashes the following commits: e3af968 [witgo] Merge branch 'master' of https://github.com/apache/spark into SPARK-1693_new dc63905 [witgo] SPARK-1693: Most of the tests throw a java.lang.SecurityException when spark built for hadoop 2.3.0 , 2.4.0 (cherry picked from commit d940e4c) Signed-off-by: Patrick Wendell <[email protected]>
pdeyhim
pushed a commit
to pdeyhim/spark-1
that referenced
this pull request
Jun 25, 2014
…n s... ...park built for hadoop 2.3.0 , 2.4.0 Author: witgo <[email protected]> Closes apache#628 from witgo/SPARK-1693_new and squashes the following commits: e3af968 [witgo] Merge branch 'master' of https://github.com/apache/spark into SPARK-1693_new dc63905 [witgo] SPARK-1693: Most of the tests throw a java.lang.SecurityException when spark built for hadoop 2.3.0 , 2.4.0
bzhaoopenstack
pushed a commit
to bzhaoopenstack/spark
that referenced
this pull request
Sep 11, 2019
Flink doesn't want ARM CI jobs run on every PR before it's stable enough. Add a comment only pipeline for this kind of requirement.
rshkv
pushed a commit
to rshkv/spark
that referenced
this pull request
Feb 27, 2020
…e query (apache#628) apache#26738 apache#26749 ### What changes were proposed in this pull request? Depend on type coercion when building the replace query. This would solve an edge case where when trying to replace `NaN`s, `0`s would get replace too. ### Why are the changes needed? This Scala code snippet: ``` import scala.math; println(Double.NaN.toLong) ``` returns `0` which is problematic as if you run the following Spark code, `0`s get replaced as well: ``` >>> df = spark.createDataFrame([(1.0, 0), (0.0, 3), (float('nan'), 0)], ("index", "value")) >>> df.show() +-----+-----+ |index|value| +-----+-----+ | 1.0| 0| | 0.0| 3| | NaN| 0| +-----+-----+ >>> df.replace(float('nan'), 2).show() +-----+-----+ |index|value| +-----+-----+ | 1.0| 2| | 0.0| 3| | 2.0| 2| +-----+-----+ ``` ### Does this PR introduce any user-facing change? Yes, after the PR, running the same above code snippet returns the correct expected results: ``` >>> df = spark.createDataFrame([(1.0, 0), (0.0, 3), (float('nan'), 0)], ("index", "value")) >>> df.show() +-----+-----+ |index|value| +-----+-----+ | 1.0| 0| | 0.0| 3| | NaN| 0| +-----+-----+ >>> df.replace(float('nan'), 2).show() +-----+-----+ |index|value| +-----+-----+ | 1.0| 0| | 0.0| 3| | 2.0| 0| +-----+-----+ ``` And additionally, query results are changed as a result of the change in depending on scala's type coercion rules. ### How was this patch tested? <!-- If tests were added, say they were added here. Please make sure to add some test cases that check the changes thoroughly including negative and positive cases if possible. If it was tested in a way different from regular unit tests, please clarify how you tested step by step, ideally copy and paste-able, so that other reviewers can test and check, and descendants can verify in the future. If tests were not added, please describe why they were not added and/or why it was difficult to add. --> Added unit tests to verify replacing `NaN` only affects columns of type `Float` and `Double`.
agirish
pushed a commit
to HPEEzmeral/apache-spark
that referenced
this pull request
May 5, 2022
* Adding SQL API to write to kafka from Spark (apache#567) * Branch 2.4.3 extended kafka and examples (apache#569) * The v2 API is in its own package - the v2 api is in a different package - the old functionality is available in a separated package * v2 API examples - All the examples are using the newest API. - I have removed the old examples since they are not relevant any more and the same functionality is shown in the new examples usin the new API. * Adding easy access to commitable offsets * Adding easy access to commitable offsets Co-authored-by: Nicolas A Perez <[email protected]>
udaynpusa
pushed a commit
to mapr/spark
that referenced
this pull request
Jan 30, 2024
* Adding SQL API to write to kafka from Spark (apache#567) * Branch 2.4.3 extended kafka and examples (apache#569) * The v2 API is in its own package - the v2 api is in a different package - the old functionality is available in a separated package * v2 API examples - All the examples are using the newest API. - I have removed the old examples since they are not relevant any more and the same functionality is shown in the new examples usin the new API. * Adding easy access to commitable offsets * Adding easy access to commitable offsets Co-authored-by: Nicolas A Perez <[email protected]>
MaxGekk
added a commit
to MaxGekk/spark
that referenced
this pull request
Oct 5, 2026
…e#628) ### What changes were proposed in this pull request? VARKA-250 (item 74.3): `VarkaBodyEmitter.emitBody` is split by role. **Before.** One 379-line method emitted all three method roles of an emitted class: the driver, a group's loop method and its epilogue, for both the dense and masked bodies, and for both driver forms. It was a run of numbered prologue steps followed by a switch on the role: 1. the empty-batch return; 2. the sizes and the scratch segment; 3. the output segments, with the driver's validity zero; 4. the inputs' null state, then the bitmap pass or the table driver's one call; 5. the all-null shortcut, then the species and literals. Every step decided for itself which role it was in: eight checks of the mode and six of `driverOutputTable`. So the table driver threaded through the unrolled driver's steps as guards. **After.** - **`emitDriver`** emits the driver: - **`emitTableDriver`**: the empty-batch return, the output-plan call, the table shortcut and the calls to the groups. It maps no segment, sizes nothing and hoists nothing, and calls no prologue helper. `Slots.plan` confirms it plans no inputs, outputs or literals for it. - **`emitUnrolledDriver`**: the prologue over every output with the validity zero or fill, the bitmap pass and unrolled shortcut in the masked driver, then the calls. - **`emitGroupBody`** emits a loop method or an epilogue: the prologue over the body's outputs, then the vector loop or the single masked pass, then the status return. - **Shared prologue steps** (`emitEmptyReturn`, `emitSizes`, `emitScratch`, `emitOutputSegments`, `emitInputState`, `emitSpecies`, `emitLiterals`) and the driver's calls (`emitGroupCalls`) are private helpers, emitted in exactly the order they always were. `VarkaLoopEmitter` calls the two entry points. The group body no longer takes the class descriptor, since only the driver's calls read it. The longest new method is 55 lines. ### Why are the changes needed? Item 74.3 of the refactoring list (`m8/SCOPE.md`). Each role's code now reads as that role's, and later emitter work can change one role without reasoning about the other two. ### Does this PR introduce _any_ user-facing change? No. No emitted byte moves under any option. ### How was this patch tested? - **`emitted_bytes.json` unchanged.** `VarkaEmittedBytesSuite` passes as it is. That oracle pins only the defaults and the `useAVX` arms, so it doesn't reach the unrolled driver, the stage driver or the budget-off form this split rewrites. - **Side by side, during development and then deleted.** The old `emitBody` was kept beside the new emitters behind a test switch. The two emitted identical bytes on all 164 arm-width pairs: the defaults and every option arm of `VarkaEmitOption.TABLE`, at both widths, over the coverage rows and 10,000 shapes from each fuzzer. With one instruction of the new driver changed, the comparison failed on the defaults at both widths, so it does bite. It ran on 20 threads in 5.5 minutes. - **`dev/varka_gate.sh`** passed every step: compile, wide, narrow (128-bit), sweep, doc, bench, lint, quotes. ### Was this patch authored or co-authored using generative AI tooling? Generated-by: Claude Code (Claude Opus 5.5)
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
...park built for hadoop 2.3.0 , 2.4.0