Repository navigation
Conversation
|
Can one of the admins verify this patch? |
There was a problem hiding this comment.
If using numPartitions = partitions, there is a chance that p1 == p2 && p1.numPartitions != p2.numPartitions is true. For example, if rdd.sample is empty, p1 = new RangePartitioner[...](10, rdd, true), and p2 = new RangePartitioner[...](1, rdd, true).
That's confusing. So I changed partitions to rangeBounds.length + 1.
|
Is there any further suggestion about this one? |
…pache#549. remove actorToWorker in master.scala, which is actually not used actorToWorker is actually not used in the code....just remove it Author: CodingCat <[email protected]> == Merge branch commits == commit 52656c2d4bbf9abcd8bef65d454badb9cb14a32c Author: CodingCat <[email protected]> Date: Thu Feb 6 00:28:26 2014 -0500 remove actorToWorker in master.scala, which is actually not used
|
Jenkins, test this please. |
|
This looks good to me pending test passes. |
|
Merged build triggered. |
|
Merged build started. |
|
Merged build finished. All automated tests passed. |
|
All automated tests passed. |
|
Thanks. I'm merging this in master. |
Adding a paragraph clarifying a weird behavior in RangePartitioner. See also #549. Author: Reynold Xin <[email protected]> Closes #1012 from rxin/partitioner-doc and squashes the following commits: 6f0109e [Reynold Xin] SPARK-1628 follow up: Improve RangePartitioner's documentation.
JIRA: https://issues.apache.org/jira/browse/SPARK-1628 Added `hashCode` in HashPartitioner, RangePartitioner, PythonPartitioner and PageRankUtils.CustomPartitioner. Author: zsxwing <[email protected]> Closes apache#549 from zsxwing/SPARK-1628 and squashes the following commits: 2620936 [zsxwing] SPARK-1628: Add missing hashCode methods in Partitioner subclasses
Adding a paragraph clarifying a weird behavior in RangePartitioner. See also apache#549. Author: Reynold Xin <[email protected]> Closes apache#1012 from rxin/partitioner-doc and squashes the following commits: 6f0109e [Reynold Xin] SPARK-1628 follow up: Improve RangePartitioner's documentation.
JIRA: https://issues.apache.org/jira/browse/SPARK-1628 Added `hashCode` in HashPartitioner, RangePartitioner, PythonPartitioner and PageRankUtils.CustomPartitioner. Author: zsxwing <[email protected]> Closes apache#549 from zsxwing/SPARK-1628 and squashes the following commits: 2620936 [zsxwing] SPARK-1628: Add missing hashCode methods in Partitioner subclasses
Adding a paragraph clarifying a weird behavior in RangePartitioner. See also apache#549. Author: Reynold Xin <[email protected]> Closes apache#1012 from rxin/partitioner-doc and squashes the following commits: 6f0109e [Reynold Xin] SPARK-1628 follow up: Improve RangePartitioner's documentation.
should use -a in tee
### What changes were proposed in this pull request?
Task 233's four laptop benchmarks, the ones post 3 ("When Spark stops compiling your query") quotes beside the runners, ran unpinned: plain `build/sbt runMain`, so the benchmark thread could move between the laptop's four Zen 5 cores at 5.16 GHz and its eight Zen 5c cores at 3.29 GHz. In the owner's idle hour they ran again on the quiet laptop, pinned to the fast cores the way `dev/varka_bench_regen.sh` pins them (`taskset -c 0-3,12-15`), and were repeated where the post quotes them: the fallback-cost and interpreter benchmarks five times, the wide projection twice (its first pinned run moved under 3% on every row), and the compile wait once. Each `*-jdk25-laptop-pinned-results.txt` holds every run as the benchmark printed it, with its provenance, including that on eight processors the JVM starts 4 C2 compiler threads instead of 12, and which runs had other work on the other core complex.
`PLAN_TASK_233.md` 17 scores every laptop number the post quotes against them, by a rule set before the repeats: a published number is corrected only where all five pinned runs fall outside it.
- The wide-projection table stands: 1.16 and 1.25 times slower for 50 and 99 cheap columns against 1.18 to 1.20 and 1.24 to 1.27 pinned, 12% and 16% faster for mixed columns against 11 to 17% and 16 to 18%. So does the time C2's code arrives in Figure 3 (13 s; 13.0 to 14.4 s pinned).
- The 150-column stage after C2 is corrected. The post said it stays 1.06 times the time outside a stage on the laptop; the five pinned runs give 0.74 to 1.00. The stage's own time is steady (406 to 423 ms); the out-of-stage time is what moves, 407 to 564 ms between runs and, within run 1, 385 in its table against 564 in its series, a JIT outcome per JVM rather than the pin. The post now says "on the laptop, 0.7 to 1.1 times, depending on the run", which covers all six runs, and Figure 3's laptop line, one of them, stays.
- The interpreter range is corrected: the post said 3.6 to 4.4 times the compiled projection on the laptop; the pinned runs give 3.06 to 3.44 at 100 columns and 3.61 to 4.20 at 1000, below the unpinned run at every width. The post now says 3.1 to 4.2.
Row 233 notes the correction. **Built on apache#544; merge apache#544 first.** **Republished** on the owner's go: the page rebuilt from this branch is live as `708fd21` in `vecbricks/vecbricks.github.io`, and the last commit here records it.
### Why are the changes needed?
A laptop number in a published post should not depend on which cores the scheduler happened to pick. The rerun was deferred to the owner's next idle-machine window, which came today.
### Does this PR introduce _any_ user-facing change?
No. Results files, the plan, and two sentences of a post.
### How was this patch tested?
Each benchmark ran one at a time with the load average under 1.0 at its start (recorded per run). `dev/varka_quote_check.py` reports zero orphans; `dev/varka_precommit.sh` reports no findings. The page rendered from this branch with `dev/varka_post_page.py` differs from the live one only in the two sentences.
### Was this patch authored or co-authored using generative AI tooling?
Generated-by: Claude Code (Claude Opus 5.5)
JIRA: https://issues.apache.org/jira/browse/SPARK-1628
Added
hashCodein HashPartitioner, RangePartitioner, PythonPartitioner and PageRankUtils.CustomPartitioner.