Skip to content

SPARK-1693: Most of the tests throw a java.lang.SecurityException when s... - #628

Closed
witgo wants to merge 2 commits into
apache:masterfrom
witgo:SPARK-1693_new
Closed

witgo wants to merge 2 commits into
apache:masterfrom
witgo:SPARK-1693_new

Conversation

@witgo

@witgo witgo commented May 4, 2014

Copy link
Copy Markdown
Contributor

...park built for hadoop 2.3.0 , 2.4.0

@AmplabJenkins

Copy link
Copy Markdown

Can one of the admins verify this patch?

@pwendell

pwendell commented May 4, 2014

Copy link
Copy Markdown
Contributor

Jenkins, test this please.

@AmplabJenkins

Copy link
Copy Markdown

Merged build triggered.

@AmplabJenkins

Copy link
Copy Markdown

Merged build started.

@AmplabJenkins

Copy link
Copy Markdown

Merged build finished. All automated tests passed.

@AmplabJenkins

Copy link
Copy Markdown

All automated tests passed.
Refer to this link for build results: https://amplab.cs.berkeley.edu/jenkins/job/SparkPullRequestBuilder/14640/

@pwendell

pwendell commented May 5, 2014

Copy link
Copy Markdown
Contributor

Thanks - I tested this locally.

asfgit pushed a commit that referenced this pull request May 5, 2014
…n s...

...park built for hadoop 2.3.0 , 2.4.0

Author: witgo <[email protected]>

Closes #628 from witgo/SPARK-1693_new and squashes the following commits:

e3af968 [witgo] Merge branch 'master' of https://github.com/apache/spark into SPARK-1693_new
dc63905 [witgo] SPARK-1693: Most of the tests throw a java.lang.SecurityException when spark built for hadoop 2.3.0 , 2.4.0
(cherry picked from commit d940e4c)

Signed-off-by: Patrick Wendell <[email protected]>
@asfgit asfgit closed this in d940e4c May 5, 2014
@witgo
witgo deleted the SPARK-1693_new branch May 5, 2014 01:36
pdeyhim pushed a commit to pdeyhim/spark-1 that referenced this pull request Jun 25, 2014
…n s...

...park built for hadoop 2.3.0 , 2.4.0

Author: witgo <[email protected]>

Closes apache#628 from witgo/SPARK-1693_new and squashes the following commits:

e3af968 [witgo] Merge branch 'master' of https://github.com/apache/spark into SPARK-1693_new
dc63905 [witgo] SPARK-1693: Most of the tests throw a java.lang.SecurityException when spark built for hadoop 2.3.0 , 2.4.0
bzhaoopenstack pushed a commit to bzhaoopenstack/spark that referenced this pull request Sep 11, 2019
Flink doesn't want ARM CI jobs run on every PR before it's stable
enough. Add a comment only pipeline for this kind of requirement.
rshkv pushed a commit to rshkv/spark that referenced this pull request Feb 27, 2020
…e query (apache#628)

apache#26738
apache#26749

### What changes were proposed in this pull request?
Depend on type coercion when building the replace query. This would solve an edge case where when trying to replace `NaN`s, `0`s would get replace too.

### Why are the changes needed?
This Scala code snippet:
```
import scala.math;

println(Double.NaN.toLong)
```
returns `0` which is problematic as if you run the following Spark code, `0`s get replaced as well:
```
>>> df = spark.createDataFrame([(1.0, 0), (0.0, 3), (float('nan'), 0)], ("index", "value"))
>>> df.show()
+-----+-----+
|index|value|
+-----+-----+
|  1.0|    0|
|  0.0|    3|
|  NaN|    0|
+-----+-----+
>>> df.replace(float('nan'), 2).show()
+-----+-----+
|index|value|
+-----+-----+
|  1.0|    2|
|  0.0|    3|
|  2.0|    2|
+-----+-----+ 
```

### Does this PR introduce any user-facing change?
Yes, after the PR, running the same above code snippet returns the correct expected results:
```
>>> df = spark.createDataFrame([(1.0, 0), (0.0, 3), (float('nan'), 0)], ("index", "value"))
>>> df.show()
+-----+-----+
|index|value|
+-----+-----+
|  1.0|    0|
|  0.0|    3|
|  NaN|    0|
+-----+-----+

>>> df.replace(float('nan'), 2).show()
+-----+-----+
|index|value|
+-----+-----+
|  1.0|    0|
|  0.0|    3|
|  2.0|    0|
+-----+-----+
```
And additionally, query results are changed as a result of the change in depending on scala's type coercion rules.

### How was this patch tested?
<!--
If tests were added, say they were added here. Please make sure to add some test cases that check the changes thoroughly including negative and positive cases if possible.
If it was tested in a way different from regular unit tests, please clarify how you tested step by step, ideally copy and paste-able, so that other reviewers can test and check, and descendants can verify in the future.
If tests were not added, please describe why they were not added and/or why it was difficult to add.
-->
Added unit tests to verify replacing `NaN` only affects columns of type `Float` and `Double`.
agirish pushed a commit to HPEEzmeral/apache-spark that referenced this pull request May 5, 2022
* Adding SQL API to write to kafka from Spark (apache#567)

* Branch 2.4.3 extended kafka and examples (apache#569)

* The v2 API is in its own package

- the v2 api is in a different package
- the old functionality is available in a separated package

* v2 API examples

- All the examples are using the newest API.
- I have removed the old examples since they are not relevant any more and the same functionality is shown in the new examples usin the new API.

* Adding easy access to commitable offsets

* Adding easy access to commitable offsets

Co-authored-by: Nicolas A Perez <[email protected]>
udaynpusa pushed a commit to mapr/spark that referenced this pull request Jan 30, 2024
* Adding SQL API to write to kafka from Spark (apache#567)

* Branch 2.4.3 extended kafka and examples (apache#569)

* The v2 API is in its own package

- the v2 api is in a different package
- the old functionality is available in a separated package

* v2 API examples

- All the examples are using the newest API.
- I have removed the old examples since they are not relevant any more and the same functionality is shown in the new examples usin the new API.

* Adding easy access to commitable offsets

* Adding easy access to commitable offsets

Co-authored-by: Nicolas A Perez <[email protected]>
MaxGekk added a commit to MaxGekk/spark that referenced this pull request Oct 5, 2026
…e#628)

### What changes were proposed in this pull request?

VARKA-250 (item 74.3): `VarkaBodyEmitter.emitBody` is split by role.

**Before.** One 379-line method emitted all three method roles of an emitted class: the driver, a group's loop method and its epilogue, for both the dense and masked bodies, and for both driver forms. It was a run of numbered prologue steps followed by a switch on the role:
1. the empty-batch return;
2. the sizes and the scratch segment;
3. the output segments, with the driver's validity zero;
4. the inputs' null state, then the bitmap pass or the table driver's one call;
5. the all-null shortcut, then the species and literals.

Every step decided for itself which role it was in: eight checks of the mode and six of `driverOutputTable`. So the table driver threaded through the unrolled driver's steps as guards.

**After.**
- **`emitDriver`** emits the driver:
  - **`emitTableDriver`**: the empty-batch return, the output-plan call, the table shortcut and the calls to the groups. It maps no segment, sizes nothing and hoists nothing, and calls no prologue helper. `Slots.plan` confirms it plans no inputs, outputs or literals for it.
  - **`emitUnrolledDriver`**: the prologue over every output with the validity zero or fill, the bitmap pass and unrolled shortcut in the masked driver, then the calls.
- **`emitGroupBody`** emits a loop method or an epilogue: the prologue over the body's outputs, then the vector loop or the single masked pass, then the status return.
- **Shared prologue steps** (`emitEmptyReturn`, `emitSizes`, `emitScratch`, `emitOutputSegments`, `emitInputState`, `emitSpecies`, `emitLiterals`) and the driver's calls (`emitGroupCalls`) are private helpers, emitted in exactly the order they always were.

`VarkaLoopEmitter` calls the two entry points. The group body no longer takes the class descriptor, since only the driver's calls read it. The longest new method is 55 lines.

### Why are the changes needed?

Item 74.3 of the refactoring list (`m8/SCOPE.md`). Each role's code now reads as that role's, and later emitter work can change one role without reasoning about the other two.

### Does this PR introduce _any_ user-facing change?

No. No emitted byte moves under any option.

### How was this patch tested?

- **`emitted_bytes.json` unchanged.** `VarkaEmittedBytesSuite` passes as it is. That oracle pins only the defaults and the `useAVX` arms, so it doesn't reach the unrolled driver, the stage driver or the budget-off form this split rewrites.
- **Side by side, during development and then deleted.** The old `emitBody` was kept beside the new emitters behind a test switch. The two emitted identical bytes on all 164 arm-width pairs: the defaults and every option arm of `VarkaEmitOption.TABLE`, at both widths, over the coverage rows and 10,000 shapes from each fuzzer. With one instruction of the new driver changed, the comparison failed on the defaults at both widths, so it does bite. It ran on 20 threads in 5.5 minutes.
- **`dev/varka_gate.sh`** passed every step: compile, wide, narrow (128-bit), sweep, doc, bench, lint, quotes.

### Was this patch authored or co-authored using generative AI tooling?

Generated-by: Claude Code (Claude Opus 5.5)
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants