Repository navigation
Cassandra support - #1452
Cassandra support#1452
Conversation
| .mapPartitions { partition: Iterator[Seq[(Long, Long)]] => | ||
| val session = instance.session | ||
|
|
||
| val statement = session.prepare(s"SELECT value FROM ${instance.keySpace}.${table} WHERE key = ?") |
There was a problem hiding this comment.
probably that would be faster to use range queries there, but not sure (the only way to figure it out -- to try)
…e store (unit by default); 2. fixed session close on read and write not breaking async and / or lazy evaluations; 3. a bit difficult session management for client code, though as a good point this management not breaks speed.
|
Hey @allixender, would you be able to review this? We're also having issues getting the unit tests to run on Travis, any thoughts? Thanks! |
|
Hi @lossyrob , @pomadchin , |
|
@allixender First of all you can look diff of this pr. Second point: we have separate sub projects for s3, accumulo, and cassandra backends, hadoop and file are included into spark sub project. What's different from your pr: we don't use datastax connector, but we use datastax java driver. Main differences in rdd reader and rdd writer So possible questions are:
By the way, cassandra space time tests are much slower than others, but it's understandable: other backends has mock clients (accumulo / s3), but for cassandra we need a real / embedded cassandra, probably that's one of the reasons why cassandra tests are all in all a bit slower. P.S. I tried to load LC8 tiles, and moved chatta demo to use any backend (including cassandra) and there were no significant differences for tile read speed / tile ingest (but i hadn't checked exact timings), and everything works as expected. |
|
Includes fix for open jdk 7: travis-ci/travis-ci#5227 |
|
Rollback to an old version and added links to travis issues. |
|
@pomadchin when I try to run the tests I get these types of errors: DeferredAbortedSuite:
[info] Exception encountered when attempting to run a suite with class name: org.scalatest.DeferredAbortedSuite *** ABORTED ***
[info] com.datastax.driver.core.exceptions.NoHostAvailableException: All host(s) tried for query failed (tried: /127.0.0.1:9042 (com.datastax.driver.core.exceptions.TransportException: [/127.0.0.1] Cannot connect))
[info] at com.datastax.driver.core.ControlConnection.reconnectInternal(ControlConnection.java:231)
[info] at com.datastax.driver.core.ControlConnection.connect(ControlConnection.java:77)
[info] at com.datastax.driver.core.Cluster$Manager.init(Cluster.java:1414)
[info] at com.datastax.driver.core.Cluster.init(Cluster.java:162)
[info] at com.datastax.driver.core.Cluster.connectAsync(Cluster.java:333)
[info] at com.datastax.driver.core.Cluster.connectAsync(Cluster.java:308)
[info] at com.datastax.driver.core.Cluster.connect(Cluster.java:250)
[info] at geotrellis.spark.io.cassandra.CassandraInstance$class.getSession(CassandraInstance.scala:18)
[info] at geotrellis.spark.io.cassandra.BaseCassandraInstance.getSession(CassandraInstance.scala:47)
[info] at geotrellis.spark.io.cassandra.CassandraInstance$class.withSessionDo(CassandraInstance.scala:34)What's with that? |
|
@lossyrob you just have no running cassandra instance on your machine https://github.com/pomadchin/geotrellis/blob/dcf5f45ef985939a0a8135926eff48473fd1493e/cassandra/src/test/scala/geotrellis/spark/io/cassandra/CassandraSpaceTimeSpec.scala#L18-L24 Use https://github.com/pomadchin/geotrellis/blob/dcf5f45ef985939a0a8135926eff48473fd1493e/scripts/cassandraTestDB.sh to start cassandra in docker |
|
@pomadchin can we have the unit tests fail in a more specific way when there is no environment? That informs users how to make them pass? This is my worry: There is a new user that wants to work with GeoTrellis. As I often do, I say, step 1 is to fork the repo, clone your fork, add upstream, and then run the unit tests. So they Is there a way we can capture the scenario of not-having-started-up-cassandra-docker-yet and explain that to users that are simply running all unit tests? |
|
@lossyrob think yes, we can add special exception type, and to check cassandra availability before all tests run. |
|
@pomadchin also we need the script to be inside the geotrellis project. It should be a one liner to run the cassandra instance for the unit tests; it should also say in a README.md in the cassandra subproject about why and how to run the cassandra docker container and the unit tests. |
|
@pomadchin nevermind, I thought you had linked to a gist. Disregard the script request. |
|
What's with the 00:14:50 DCAwareRoundRobinPolicy: Using data-center name 'datacenter1' for DCAwareRoundRobinPolicy (if this is incorrect, please provide the correct datacenter name with DCAwareRoundRobinPolicy constructor)a bunch of times in the last build failure logs? |
|
It is related to a load balancing policy, would add configuration settings, forgot about it. However these annoying messages still would be present, nothing to do with it, it's |
No description provided.