Skip to content

SPARK-1628: Add missing hashCode methods in Partitioner subclasses - #549

Closed
zsxwing wants to merge 1 commit into
apache:masterfrom
zsxwing:SPARK-1628
Closed

zsxwing wants to merge 1 commit into
apache:masterfrom
zsxwing:SPARK-1628

Conversation

@zsxwing

@zsxwing zsxwing commented Apr 25, 2014

Copy link
Copy Markdown
Member

JIRA: https://issues.apache.org/jira/browse/SPARK-1628

Added hashCode in HashPartitioner, RangePartitioner, PythonPartitioner and PageRankUtils.CustomPartitioner.

@AmplabJenkins

Copy link
Copy Markdown

Can one of the admins verify this patch?

Copy link
Copy Markdown
Member Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

If using numPartitions = partitions, there is a chance that p1 == p2 && p1.numPartitions != p2.numPartitions is true. For example, if rdd.sample is empty, p1 = new RangePartitioner[...](10, rdd, true), and p2 = new RangePartitioner[...](1, rdd, true).

That's confusing. So I changed partitions to rangeBounds.length + 1.

@zsxwing

zsxwing commented May 4, 2014

Copy link
Copy Markdown
Member Author

Is there any further suggestion about this one?

pwendell pushed a commit to pwendell/spark that referenced this pull request May 12, 2014
…pache#549.

remove actorToWorker in master.scala, which is actually not used

actorToWorker is actually not used in the code....just remove it

Author: CodingCat <[email protected]>

== Merge branch commits ==

commit 52656c2d4bbf9abcd8bef65d454badb9cb14a32c
Author: CodingCat <[email protected]>
Date:   Thu Feb 6 00:28:26 2014 -0500

    remove actorToWorker in master.scala, which is actually not used
@rxin

rxin commented Jun 8, 2014

Copy link
Copy Markdown
Contributor

Jenkins, test this please.

@rxin

rxin commented Jun 8, 2014

Copy link
Copy Markdown
Contributor

This looks good to me pending test passes.

@AmplabJenkins

Copy link
Copy Markdown

Merged build triggered.

@AmplabJenkins

Copy link
Copy Markdown

Merged build started.

@AmplabJenkins

Copy link
Copy Markdown

Merged build finished. All automated tests passed.

@AmplabJenkins

Copy link
Copy Markdown

All automated tests passed.
Refer to this link for build results: https://amplab.cs.berkeley.edu/jenkins/job/SparkPullRequestBuilder/15541/

@rxin

rxin commented Jun 8, 2014

Copy link
Copy Markdown
Contributor

Thanks. I'm merging this in master.

@asfgit asfgit closed this in a71c6d1 Jun 8, 2014
asfgit pushed a commit that referenced this pull request Jun 9, 2014
Adding a paragraph clarifying a weird behavior in RangePartitioner.

See also #549.

Author: Reynold Xin <[email protected]>

Closes #1012 from rxin/partitioner-doc and squashes the following commits:

6f0109e [Reynold Xin] SPARK-1628 follow up: Improve RangePartitioner's documentation.
@zsxwing
zsxwing deleted the SPARK-1628 branch June 19, 2014 14:33
pdeyhim pushed a commit to pdeyhim/spark-1 that referenced this pull request Jun 25, 2014
JIRA: https://issues.apache.org/jira/browse/SPARK-1628

Added `hashCode` in HashPartitioner, RangePartitioner, PythonPartitioner and PageRankUtils.CustomPartitioner.

Author: zsxwing <[email protected]>

Closes apache#549 from zsxwing/SPARK-1628 and squashes the following commits:

2620936 [zsxwing] SPARK-1628: Add missing hashCode methods in Partitioner subclasses
pdeyhim pushed a commit to pdeyhim/spark-1 that referenced this pull request Jun 25, 2014
Adding a paragraph clarifying a weird behavior in RangePartitioner.

See also apache#549.

Author: Reynold Xin <[email protected]>

Closes apache#1012 from rxin/partitioner-doc and squashes the following commits:

6f0109e [Reynold Xin] SPARK-1628 follow up: Improve RangePartitioner's documentation.
xiliu82 pushed a commit to xiliu82/spark that referenced this pull request Sep 4, 2014
JIRA: https://issues.apache.org/jira/browse/SPARK-1628

Added `hashCode` in HashPartitioner, RangePartitioner, PythonPartitioner and PageRankUtils.CustomPartitioner.

Author: zsxwing <[email protected]>

Closes apache#549 from zsxwing/SPARK-1628 and squashes the following commits:

2620936 [zsxwing] SPARK-1628: Add missing hashCode methods in Partitioner subclasses
xiliu82 pushed a commit to xiliu82/spark that referenced this pull request Sep 4, 2014
Adding a paragraph clarifying a weird behavior in RangePartitioner.

See also apache#549.

Author: Reynold Xin <[email protected]>

Closes apache#1012 from rxin/partitioner-doc and squashes the following commits:

6f0109e [Reynold Xin] SPARK-1628 follow up: Improve RangePartitioner's documentation.
bzhaoopenstack pushed a commit to bzhaoopenstack/spark that referenced this pull request Sep 11, 2019
MaxGekk added a commit to MaxGekk/spark that referenced this pull request Oct 2, 2026
### What changes were proposed in this pull request?

Task 233's four laptop benchmarks, the ones post 3 ("When Spark stops compiling your query") quotes beside the runners, ran unpinned: plain `build/sbt runMain`, so the benchmark thread could move between the laptop's four Zen 5 cores at 5.16 GHz and its eight Zen 5c cores at 3.29 GHz. In the owner's idle hour they ran again on the quiet laptop, pinned to the fast cores the way `dev/varka_bench_regen.sh` pins them (`taskset -c 0-3,12-15`), and were repeated where the post quotes them: the fallback-cost and interpreter benchmarks five times, the wide projection twice (its first pinned run moved under 3% on every row), and the compile wait once. Each `*-jdk25-laptop-pinned-results.txt` holds every run as the benchmark printed it, with its provenance, including that on eight processors the JVM starts 4 C2 compiler threads instead of 12, and which runs had other work on the other core complex.

`PLAN_TASK_233.md` 17 scores every laptop number the post quotes against them, by a rule set before the repeats: a published number is corrected only where all five pinned runs fall outside it.

- The wide-projection table stands: 1.16 and 1.25 times slower for 50 and 99 cheap columns against 1.18 to 1.20 and 1.24 to 1.27 pinned, 12% and 16% faster for mixed columns against 11 to 17% and 16 to 18%. So does the time C2's code arrives in Figure 3 (13 s; 13.0 to 14.4 s pinned).
- The 150-column stage after C2 is corrected. The post said it stays 1.06 times the time outside a stage on the laptop; the five pinned runs give 0.74 to 1.00. The stage's own time is steady (406 to 423 ms); the out-of-stage time is what moves, 407 to 564 ms between runs and, within run 1, 385 in its table against 564 in its series, a JIT outcome per JVM rather than the pin. The post now says "on the laptop, 0.7 to 1.1 times, depending on the run", which covers all six runs, and Figure 3's laptop line, one of them, stays.
- The interpreter range is corrected: the post said 3.6 to 4.4 times the compiled projection on the laptop; the pinned runs give 3.06 to 3.44 at 100 columns and 3.61 to 4.20 at 1000, below the unpinned run at every width. The post now says 3.1 to 4.2.

Row 233 notes the correction. **Built on apache#544; merge apache#544 first.** **Republished** on the owner's go: the page rebuilt from this branch is live as `708fd21` in `vecbricks/vecbricks.github.io`, and the last commit here records it.

### Why are the changes needed?

A laptop number in a published post should not depend on which cores the scheduler happened to pick. The rerun was deferred to the owner's next idle-machine window, which came today.

### Does this PR introduce _any_ user-facing change?

No. Results files, the plan, and two sentences of a post.

### How was this patch tested?

Each benchmark ran one at a time with the load average under 1.0 at its start (recorded per run). `dev/varka_quote_check.py` reports zero orphans; `dev/varka_precommit.sh` reports no findings. The page rendered from this branch with `dev/varka_post_page.py` differs from the live one only in the two sentences.

### Was this patch authored or co-authored using generative AI tooling?

Generated-by: Claude Code (Claude Opus 5.5)
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

4 participants